Five diffusion papers from the September 29 batch: a belief state on the simplex, critic error projected out of the update, and guidance held inside its subspace

Five diffusion papers from the September 29 batch: a belief state on the simplex, critic error projected out of the update, and guidance held inside its subspace

A ranked scan of five diffusion-model preprints from the September 29, 2026 arXiv batch: discrete diffusion that carries a belief over categories instead of sampling it away, one-step generation trained by matching mixtures of features, one-line removal of the critic error that degrades distribution-matching distillation, a training regularizer that keeps mixture-of-experts guidance inside the conditioned subspace, and reinforcement learning that matches trajectory distributions instead of maximizing reward.

arXiv's newest announcement is the batch for Tuesday, 29 September 2026: 460 new cs.CV submissions and 645 new cs.LG submissions. 12 A listing page groups everything arXiv announced together, and its section headers say nothing about when any single entry was written, so every candidate was checked against the arXiv API one identifier at a time. All five papers below carry verified v1 timestamps on 28 September 2026 (UTC), the earliest at 07:59 and the latest at 17:59.
The announced batch is the unit arXiv actually publishes, and it closes at arXiv's own submission cutoff rather than at this digest's publish time. Measured as a strict 24-hour window back from today's 09:00 (-05:00) publication time, three of the five carry timestamps inside the window and two fall several hours before it. Each paper's verified v1 date appears in its entry.
Citation momentum still cannot rank anything at this age, and three routes were tried. An OpenAlex lookup by arXiv DOI returned HTTP 404 on all five identifiers. 3 Semantic Scholar's graph API returned HTTP 429 on three of them and HTTP 404 on the other two. 4 A third route, tried for the first time here, is the DOI registration record: DataCite answered for all five with a citation count of zero and zero views and downloads on each, which is what a preprint submitted yesterday looks like rather than a ranking signal. 5 The ordering below therefore rests on method novelty, author-reported evidence, disclosed affiliation and public inspectability. Every number in this digest is author-reported unless a sentence says otherwise.

The five at a glance

RankPaperCentral moveStrongest reported signalInspectability
1Simplex Diffusion ModelsCarrying a belief over categories on the probability simplex through every denoising step17.0 GenPPL at 5.46 unigram entropy on OpenWebText in 64 steps, against 5.46 nats for real validation data 6Paper and pseudocode only
2MGFlowModelling the feature distribution of one-step generation as a Gaussian mixtureFDr⁶ of 1.45 on pMF-H and 1.64 on JiT-H at ImageNet 256×256, 23% and 38% below FD-Loss 7Project page with samples
3PDMDRemoving the critic's error direction from the distillation updateVBench total 83.73 at 4 NFE on Wan2.1, 1.03 points above matched DMD 8Code, weights and project page released
4SAGEAligning the unconditional MoE activation into the conditional subspacePeak DPG-Bench up 9.3% and worst-case class drift cut 9.2×, at zero inference cost 9Paper only
5Uni-TMPOMatching the policy to a reward-derived target distribution instead of maximizing rewardGenEval 0.954 and PickScore 24.301 on FLUX.1-dev, ahead of Flow-GRPO's 0.946 and 24.226 10Paper only

1. Simplex Diffusion Models: keeping the belief instead of sampling it away

Team and affiliation. Justin Deschenaux at Google DeepMind and EPFL; Alexandre Galashov at Google DeepMind and UCL Gatsby; Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet and Valentin De Bortoli at Google DeepMind. arXiv v1 dated 28 September 2026. 6
What changes. A discrete diffusion model generates a sequence of categories — tokens, characters, atoms — and at every denoising step it picks one category and throws the rest away. The paper's name for the cost is information collapse: the model cannot change its mind later about a choice it committed to early, and the uncertainty that a continuous diffusion model carries through its trajectory has no place to live. Simplex Diffusion Models (SDMs) give it one. 11
Technical insight. The belief state over the vocabulary is itself a point on the probability simplex, so the whole process can be lifted there. The forward process is a Dirichlet path towards the uniform distribution, chosen so that the reverse transition has a closed form and training reduces to a cross-entropy loss. Sampling is DDIM-like, with a churn parameter κ that tunes how much stochasticity the reverse step injects, and that parameter turns out to matter: the paper evaluates κ of 0, 0.2 and 1 across step budgets rather than fixing one. An earlier proposal in this space, Dirichlet Flow Matching, needs an ODE integrated at sample time; the SDM sampler needs no such solve. 11
Two line charts titled NFE vs TinyGSM accuracy, one at temperature 0.1 and one at temperature 1.0, plotting accuracy against number of function evaluations on a logarithmic axis, with dashed reference lines at 62.6% for greedy autoregressive decoding and 52.6% for autoregressive sampling at temperature 1.0
Figure 1 of the paper. The orange markers are the SDM variants at churn κ = 1, the blue and green markers are masked and uniform discrete diffusion, and the dashed lines are autoregressive baselines. Read the height the orange series reaches by 128 evaluations against where the other families flatten. 11
Quantitative evidence. On OpenWebText at 1024 tokens, SDMs reach 17.0 generative perplexity at 5.46 unigram entropy in 64 sampling steps, which the authors place next to real validation data at an entropy of 5.46 nats. On TinyGSM code generation at temperature 0.1, SDMs score 49.0% against 45.8% for masked and uniform diffusion, and they do it without self-conditioning, which those baselines use. Distilled to eight steps, SDMs solve 32.1% of GSM8K problems against 21.4% for distilled discrete diffusion models given 128 steps. Sudoku, molecular generation on SAFE-GPT and language-understanding benchmarks are reported in the appendices. 11
Caveat. The evaluation is entirely in discrete sequence domains. The paper's strongest claims rest on TinyGSM, GSM8K, Sudoku, OpenWebText and molecular generation, so a reader working in continuous image or video diffusion is reading a framework proposal rather than a result they can adopt. The OpenWebText frontier also depends on a sequence-frequency penalty the paper applies across every model family it compares, including the autoregressive baseline, so part of the gap it reports is produced by that shared intervention. 11
Inspectability. The preprint announces no repository and no released weights. Training, inference and the accelerated Gamma sampler are given as self-contained pseudocode in the paper, and the HTML version renders the full derivations. 11
Why it matters. Discrete diffusion has been losing arguments to autoregressive models on reasoning tasks, and the explanation on offer has been that categorical sampling destroys information. This paper removes the categorical sampling and keeps the recursive loss unchanged, which makes the explanation testable. If you work on discrete diffusion for code, molecules or text, the piece to read closely is the concentration schedule — the rule that decides how fast the Dirichlet path sharpens — because that is the knob the two schedules in the paper disagree about. 11

2. MGFlow: a mixture where one-step generation used a single Gaussian

Team and affiliation. Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui and Miao Liu, from Tsinghua University, Fudan University, Xi'an Jiaotong University, Peking University, Zhejiang University, DeepSeek-AI, ByteDance Seed, the University of California, Berkeley and BAAI. Zhang, Haoyang and Liu are marked equal contribution, with correspondence to Miao Liu. arXiv v1 dated 28 September 2026. 7
What changes. One-step visual generation is trained by comparing the features of real images with the features of generated ones inside a frozen encoder, and that comparison needs a model of what each feature population looks like. The current recipe, FD-Loss, models each population as a single Gaussian and matches their means and covariances. MGFlow models it as a Gaussian mixture with a component count you choose, which lets a population with several visual modes keep them. 12
Technical insight. The paper first separates two decisions that the existing recipes bundle: the distribution model ℳ, and the discrepancy 𝒟 used to compare distributions. A Wasserstein gradient flow then connects the global objective to the per-feature update, and the known methods fall out as choices of that pair — FD-Loss is a single Gaussian with a squared Wasserstein distance, Gaussian-kernel Drifting is a kernel density estimate with a KL divergence. MGFlow takes the middle setting: a Gaussian mixture, matched by either optimal transport or a score-based criterion. The part that makes mixtures actually work is outside the model class. Mixture expressivity alone does not stop mode collapse, so the method couples a mass-constrained assignment of samples to components with updates applied component by component, which keeps every component fed. 12
A four-panel figure: on the left, the distributional-training pipeline of batch sampling, a frozen representation encoder, a distribution model and a matching discrepancy; on the right, three ways to model a two-dimensional feature cloud — a single Gaussian, a Gaussian mixture of three overlapping components, and a kernel density estimate of the samples
Figure 2 of the paper. The three right-hand panels are the same feature cloud read at three granularities, and the labels underneath are the whole argument: global moments only, component-wise moments, local kernel estimates. MGFlow is the middle panel. 12
Quantitative evidence. On ImageNet 256×256, MGFlow reaches an FDr⁶ of 1.45 on the pMF-H backbone and 1.64 on JiT-H, which the paper reports as 23% and 38% better than the FD-Loss baseline under matched sampling and encoder settings. For text-to-image it post-trains FLUX.2 [klein] 4B into a one-step generator reaching GenEval 0.900 and PickScore 21.98, above every method the paper compares against. Ablations cover component count, the choice between optimal transport and score matching, and component assignment; training-time measurements are reported per optimizer update on eight H200 GPUs. 12
Caveat. The ImageNet comparison is against FD-Loss under the paper's own encoder set, and the reference mixture is fitted to the real feature population, so the reported margin measures the distance between two distribution models rather than against an external leaderboard. The text-to-image claim compares the one-step post-trained 4B model against the same model's four-step mode, which establishes that post-training preserved quality rather than that it beats other one-step systems. 12
Inspectability. A project page at shihaoyang0423.github.io/MGFlow-website carries generated samples for the ImageNet backbones and the text-to-image model. The preprint announces no code repository and no weights. 12
Why it matters. Distributional training has become the way one-step image generators are trained, and the single-Gaussian assumption inside it is easy to inherit without noticing. This paper tells you the assumption is one point in a two-dimensional design space, and hands you a second point. For anyone already running FD-Loss, the transferable question is how many modes the encoder space of your own model actually has, since that is what sets the component count. 12

3. PDMD: taking the critic's error out of the distillation step

Team and affiliation. Zimo Wang, Junkun Yuan, Angtian Wang, Haotian Yang, Canyu Zhang, Siyuan Yuan, Xingchang Huang, Bo Liu, Yizhi Wang, Yiding Yang, Chongyang Ma and Gordon Guocheng Qian, all listed at ByteDance Inc.; the preprint marks Wang's contribution as work done during an internship at ByteDance. arXiv v1 dated 28 September 2026. 8
What changes. Distribution Matching Distillation (DMD) compresses a video generator from tens of denoising steps to a handful, and it works by moving the student along the difference between what the frozen teacher says and what an online critic says. That critic is itself being learned, so its error enters the student's updates and accumulates. The visible symptom is progressive oversaturation and texture artefacts appearing part-way through training. PDMD keeps the recipe and removes one component of each update. 13
Technical insight. Both the teacher and the student produce an endpoint prediction from the same noisy latent, and their difference r is known at every update. PDMD projects the DMD update d onto the subspace orthogonal to r, discarding the part of the update that runs parallel to the student–critic endpoint residual. The paper shows that this residual is an unbiased estimate of the critic's endpoint error, so the discarded component is the one carrying the error rather than the signal; under high-dimensional assumptions the projection strips a constant fraction of critic error while giving up a vanishing fraction of the ideal DMD update. The implementation change is one line, with no added loss, network, data pass or training stage. 13
A two-dimensional schematic: a noise point, a noised latent x_t on a dashed trajectory to a student endpoint, teacher and critic score arrows, the residual r drawn between the critic endpoint and the student endpoint, and two short arrows labelled DMD update and PDMD update leaving the latent in different directions
Figure 3 of the paper. The two arrows at the left are the comparison: the orange DMD update carries a component along the residual r, and the green PDMD update is what remains after that component is projected out. The residual is computable at every step, which is what makes the projection free. 13
Quantitative evidence. On Wan2.1-T2V-1.3B, PDMD reaches a VBench total score of 83.73 at four function evaluations, 1.03 points above its matched DMD baseline. On MiniMax-H3-33B joint video–audio generation it reaches 83.17 total on VideoGen-Eval, 0.41 points above the strongest distilled baseline, and takes the highest score on all six audio metrics among the four-step models compared. The mechanism is isolated in a two-dimensional study, where removing the projection collapses the sample to a single mode in 14 of 20 runs and keeping it does so in 3 of 20. User studies, run against several four-step baselines on visual quality, motion and audio, are reported as favouring PDMD. 13
Caveat. The headline video gains are 1.03 and 0.41 points on benchmarks whose own spread is not established by this paper, so the ordering between methods is firmer than the size of the margins. The clean causal evidence for the mechanism comes from the two-dimensional diagnostic rather than from the video runs, and the proof of the unbiasedness property rests on assumptions the paper states rather than tests at the scale of the experiments. The authors also report that the method degrades in the single-step setting, which their own limitation section shows. 13
Inspectability. Code and models are released at the project page pdmd2026.github.io, with the repository at ZeamoxWang/pdmd. 13
Why it matters. Anyone who has watched a DMD run go bad after a few hundred iterations has been reading the symptom without a rule for what to delete. PDMD offers a rule that costs one line and needs nothing new to be trained, and it is the only paper of these five that ships the weights to test it with. The number to look for in your own run is the epoch at which oversaturation used to start. 13

4. SAGE: what CFG amplifies when the two branches route differently

Team and affiliation. Boyu Zhang, Yangming Cheng, Ning Zhang, Pengfei Liu, Weijie Li, Yifan Gao, Hangyu Li and Litong Gong, all at Alibaba Token Hub, Alibaba Group, with Litong Gong marked corresponding. arXiv v1 dated 28 September 2026. 9
What changes. Mixture-of-Experts diffusion transformers are the current recipe for scaling image generation, and classifier-free guidance is the standard way to sharpen their output. The two combine badly at high guidance, and the failure had not been named. SAGE names it routing-induced subspace leakage: because the conditional and unconditional passes go through the router independently, they select different experts, and the two realized activation sets therefore sit in different subspaces. The guided velocity is a difference between those two activations, so the part that falls outside the conditional subspace is added to the sample and then multiplied by the guidance scale. 14
Technical insight. The fix works on the unconditional write during training. SAGE takes the conditional block output, computes an orthonormal basis for the subspace it spans with a QR decomposition, projects the unconditional output onto that basis, and penalizes the residual — the loss is the squared norm of the projected-out part divided by the squared norm of the unconditional output. The regularizer changes nothing about which experts are chosen. The paper reports that conditional and unconditional routing similarity tracks the unregularized baseline closely, and that per-expert load divergence is preserved, which is the check that matters for anyone worried the fix would undo the reason to use a mixture at all. Inference is unchanged. 14
A two-part figure: on top, a diagram of a mixture-of-experts layer under classifier-free guidance where the conditional branch selects experts 1 and 3 while the unconditional branch selects experts 2 and 3; below, a geometric panel contrasting standard guidance, where the guided velocity leaves the conditional subspace, with the alignment procedure of a basis, a projection and a minimized residual
Figure 1 of the paper. The top row is the diagnosis — two routers, two expert sets, two subspaces — and the bottom right is the remedy, which is a QR basis of the conditional output and a residual to shrink. The bottom left panel is what the guided velocity does without it. 14
Quantitative evidence. In a toy experiment on class-conditional generation, SAGE cuts the worst-case per-class centre drift at high guidance by 9.2× and largely restores mode purity where the baseline has collapsed. Scaled to a 1B-parameter MoE diffusion transformer on text to image, it improves the peak DPG-Bench score by 9.3% and lifts the peak GenEval score by 2.7 percentage points, with the largest per-task GenEval gain on colour-attribute binding at 11.5 points. The baseline loses roughly 10.0 points of DPG-Bench as guidance rises. The paper also reports that the obvious alternative — a KL constraint forcing the two routing distributions together — hurts every metric, because it flattens expert specialization. 14
Caveat. The scaling evidence stops at a 1B-parameter model trained on a curated 100M-image subset, so what happens to the effect at frontier scale is untested here. The quantitative gain and the qualitative gain are not the same size: 9.3% on peak DPG-Bench sits beside 2.7 points on GenEval, and the paper's own per-task breakdown shows the GenEval gain concentrated in one category rather than spread across all of them. 14
Inspectability. The preprint announces no code repository and no weights. The training algorithm with SAGE is given in full, and the HTML version carries the theoretical analysis of what the regularizer minimizes. 14
Why it matters. Two design choices that each work well on their own are now standard together, and this paper explains a failure that appears only in the combination. If you are serving or training an MoE diffusion model and you have been treating a quality cliff at high guidance as a property of guidance itself, the diagnosis here says otherwise and the remedy is a training-time term you can add to a run you are doing anyway. 14

5. Uni-TMPO: matching a trajectory distribution instead of maximizing a reward

Team and affiliation. Zhiyuan Ma, Jiaming Li, Lingzhen Li, Dingkang Liang and Jianjun Li at the School of Computer Science and Technology, Huazhong University of Science and Technology; Yu Liu at the Institute of Information Engineering, Chinese Academy of Sciences; Xuekai Zhu at Shanghai Jiao Tong University; Kaiyan Zhang at Frontis.AI; Bowen Zhou in the Department of Electronic Engineering, Tsinghua University; Xiang Bai at the School of Software Engineering, Huazhong University of Science and Technology. arXiv v1 dated 28 September 2026. 15
What changes. Reinforcement learning has become the standard way to post-train a diffusion or flow model against a reward, and the standard objective maximizes that reward. The paper documents the consequence: the policy collapses onto a single high-reward mode, and it does so even with a reference KL term or an entropy bonus in place. Uni-TMPO replaces the maximization with a distribution match — the policy is fit to a target distribution built from the reward instead of being pushed towards its maximum. 10
Technical insight. Inside each group of trajectories sampled for one prompt, the rewards are standardized and turned into a target distribution over the group, and the policy distribution is read off the trajectories' log probabilities. Optimization then minimizes the forward KL between the two, so allocating probability mass to several good trajectories is as acceptable as allocating it to one. Two sampling schemes feed the groups: progress-conditioned coarse-to-fine sampling builds text-to-image trajectory groups, and feedback-conditioned action-chunk sampling builds vision-language-action groups from a shared initialization. The same objective covers both, which is what the "unified" in the name refers to. 10
A wide framework diagram: task inputs for text-to-image and vision-language-action on the left, a central panel showing progress-conditioned tree branching and observation-conditioned action chunks under shared optimization, and on the right a radar chart of text-to-image metrics plus a bar chart of held-out task and scene gains for two policy backbones
Figure 1 of the paper. The bottom strip carries the comparison that the rest of the figure sets up: maximizing a scalar reward concentrates the group on one sample, while matching the allocation keeps the group spread across several. The radar chart and the right-hand bars are the results that follow. 10
Quantitative evidence. On FLUX.1-dev with LoRA, post-training against a compositional-correctness reward reaches GenEval 0.954 against 0.946 for Flow-GRPO and 0.647 for the unmodified model, in 91.9 seconds per iteration against Flow-GRPO's 160.8. Against a text-rendering reward it reaches OCR accuracy 0.951 against 0.944, and against a human-preference reward PickScore 24.301 against 24.226, in 68.3 seconds per iteration against 109.1. Diversity moves further than the reward: cosine diversity is 0.248 on the GenEval protocol against 0.198 for Flow-GRPO. On the robotics side, a π₀ backbone reaches 97.8 average success across the LIBERO suites and 88.6 on MetaWorld-MT50 against Flow-SDE's 78.1, and a π₀.₅ backbone reaches 89.2 five-subtask completion on CALVIN-D against Flow-SDE's 87.0. 10
Caveat. The text-to-image study uses one backbone with one adaptation method, and the reward gaps that headline it are small — 0.954 against 0.946 on a near-binary metric, 24.301 against 24.226 on a continuous one — so the diversity and iteration-time differences carry more of the argument than the reward scores do. The larger claims live on the robotics side, where held-out tasks and scene transfer are what the method is supposed to protect, and those numbers come from two simulation benchmarks. 10
Inspectability. The preprint announces no code repository and no released weights. The causal experiments it leans on include toy distributions, a WallDetour navigation task and a real-robot dual-target placement, which the paper reports as recovering the alternative strategy after the preferred one is blocked. 10
Why it matters. Reward maximization is the default, and its failure mode looks like success on the training metric: the samples are high-reward and nearly identical. This paper connects that symptom to the objective rather than to the regularization strength, and it does so inside a framework that already covers vision-language-action policies. The result worth reading in depth is the held-out comparison, since that is where a collapsed policy and a spread one should differ most. 10

What the five have in common

Each of these five methods refuses to let a distribution be reduced to a single number, and each is defined by the structure it keeps instead. Simplex Diffusion Models keeps the belief over categories rather than the sampled category. MGFlow keeps the mixture rather than one Gaussian's moments. PDMD keeps the residual direction and deletes the component of the update aligned with it. SAGE keeps the basis of the conditional subspace and penalizes whatever the unconditional write puts outside it. Uni-TMPO keeps the whole trajectory-group distribution and matches it, rather than chasing the scalar its rewards came from.
That shared shape is also what makes them comparable. In every case the structure being preserved is something you can compute during the run — a Dirichlet parameter, a component assignment, an endpoint residual, a QR basis, a per-group reward distribution — and in every case the method is a rule for acting on it. Where a paper's rule needs the structure to be recoverable at all is where its applicability to your own pipeline gets decided, and the four entries above that state a caveat about scale are stating exactly that boundary.
Three of the five sit in cs.CV and two in cs.LG. The split follows the papers: PDMD, SAGE and Uni-TMPO are about image and video generation specifically, while Simplex Diffusion Models and MGFlow are about the training objective and could be aimed at a different data type without changing their argument.
A reading order follows from what is open in your own work today:
  • Start with PDMD if you are running or debugging a distribution-matching distillation, since it is a one-line change and the weights are published.
  • Start with SAGE if you train or serve a mixture-of-experts diffusion transformer and have met a quality cliff as guidance rises.
  • Start with MGFlow if your one-step generator is trained against a frozen encoder and you have taken the single-Gaussian feature model for granted.
  • Start with Simplex Diffusion Models if you work on discrete diffusion or on the reasoning benchmarks where it has been losing to autoregressive models.
  • Start with Uni-TMPO if you post-train flow or diffusion policies with reinforcement learning and your samples have started to look alike.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content