
Five diffusion papers from the September 17 batch: noise coupling, gaze-grounded synthesis, and a physical video world model
A ranked scan of five diffusion-model preprints from the September 17, 2026 arXiv batch, covering noise-data coupling in flow matching, gaze-grounded synthetic training data, a physical video world model, 3D garment dynamics, and talking-avatar conditioning.
arXiv announced the Thursday, 17 September 2026 batch this morning: 98 new
cs.CV submissions and 108 new cs.LG submissions, covering the submission cycle that closed with arXiv's 16 September 2026 cutoff. 12 Five of those papers are diffusion-model work. All five carry verified v1 timestamps of 15–16 September 2026 (UTC); the newest, VibeAvatar, was submitted on 16 September at 13:20 UTC.Citation counts and social engagement metrics do not exist yet for preprints this fresh, so the ranking below rests on method novelty, author-reported evidence, disclosed affiliation, and public inspectability. Every number in this digest is author-reported unless the entry says otherwise.
The five at a glance
| Rank | Paper | Central move | Strongest reported signal | Inspectability |
|---|---|---|---|---|
| 1 | Contrastive Noise Alignment | Optimising the noise distribution itself during flow-matching training | CIFAR-10 FID 56.40 at 2 Euler NFEs vs 175.91 for independent coupling | Proofs and pseudocode, no code release |
| 2 | GazeDiT | Grounding a 4D gaze label in pupil/iris geometry for synthetic eye-tracking data | Downstream gaze error on difficult cases 3.05° → 2.80° | Internal VR cohort; no code release |
| 3 | StrucPhysVideo | Physics-focused corpus curation feeding a 30B sparse-MoE video world model | Physics-IQ Verified 45.5%, ahead of Cosmos3-Super-Image2Video by 2.8 points | Code and project page released |
| 4 | DiT-Garment | A 2D diffusion transformer learning 3D garment deformation in UV space | Vertex error 5.20 cm vs 6.95 (DPF) and 15.35 (ContourCraft) on unseen garments | Code and data for research |
| 5 | VibeAvatar | Splitting phonetic accuracy and motion aesthetics across two training stages | HDTF Sync-C 8.639, FID 22.033; 10-second 512px video in under 10 seconds | ACM MM 2026 paper; no code link in the preprint |
1. Contrastive Noise Alignment: teaching the noise distribution where to go
Team and affiliation. Lennart Wittke, who is affiliated with ETH Zürich and Disney Research|Studios, and Vinicius Azevedo of Disney Research|Studios. The paper is an arXiv v1 submission dated 16 September 2026. 3
What changes. Flow-matching and diffusion training corrupt each training image with an independently sampled Gaussian vector. That pairing is arbitrary: nothing connects the noise sample to the image it will be asked to reconstruct, so the network has to learn a curved transport between two unrelated endpoints. Optimal-transport variants fix half of the problem by reassigning which noise sample pairs with which image, while the noise distribution itself stays passive. This paper optimises the noise distribution during training instead. 4
Technical insight. CNA treats one batch of noise vectors as an interacting particle system and moves the particles. A cross-modal InfoNCE term attracts each noise particle toward its paired data target, while an angular entropy term and a radial norm penalty stop the batch from collapsing onto that target. The paper proves that the equilibrium of these three forces asymptotically preserves Gaussian structure, which is what keeps inference unchanged: the same network, sampler and function-evaluation count are used at test time. 4

Quantitative evidence. On unconditional CIFAR-10, FID at one Euler step falls from 346.76 with independent couplings and 229.98 with optimal-transport couplings to 148.54 for CNA combined with OT initialisation. At two steps the three figures are 175.91, 91.75 and 56.40; at four steps, 54.39, 30.00 and 22.64. On ImageNet32 at one step the corresponding numbers are 419.92, 252.88 and 166.07. The abstract summarises the few-step regime as more than 50% FID reduction against standard rectified flow and at least 24% against optimal-transport baselines, across 2–4 function evaluations. 4
Inspectability. The preprint is 21 pages with full proofs, a pseudocode listing, hyperparameter tables and curvature measurements. A code repository or project page is absent from the submission. 3
Caveat. The authors name the central tension themselves: aligning the noise prior aggressively pulls against preserving its Gaussian properties, and they are still looking for regularisers that hold the adapted prior on the Gaussian manifold. The current method starts from random noise rather than constructing the prior from data, so the starting point can be sub-optimal. Conditional and latent generation are listed as work in progress, which means the reported results cover pixel-space unconditional training only. 4
Why it matters. Anyone training a flow-matching or rectified-flow model can apply this as a change to the training objective with no architectural work, and the gain lands precisely where few-step samplers are weakest. If the curvature reduction holds outside CIFAR-10-scale images, it competes with distillation for the same budget.
2. GazeDiT: making a synthetic image carry its own label
Team and affiliation. Dongze Wu (Meta and Georgia Institute of Technology), David Colmenares, Fengting Yang, Jogendra Nath Kundu, Ali Behrooz and Conny Lu (Meta), and Yao Xie (Georgia Institute of Technology). The paper is an arXiv v1 submission dated 15 September 2026. 5
What changes. Diffusion models increasingly generate training data, but a low-dimensional continuous label is hard to control: a text prompt can be checked against a description, while eye-tracking supervision needs each image to match a numeric 4D binocular gaze. In near-eye images that gaze lives in subtle pupil and iris geometry. GazeDiT builds a spatial condition inside the pipeline so the label is expressed where it physically appears, rather than being handed to the model as a bare number. 6
Technical insight. During training, a frozen SegFormer segments pupil and iris masks from real near-eye images, and the mask latent is concatenated channel-wise with the image latent before patch embedding, with the gaze vector supplied as a token. The generator therefore learns appearance conditioned on geometry that a real segmenter measured. At inference no source image is needed: a physical eye renderer samples gaze-consistent geometry by varying eye anatomy and camera state, so the model can synthesise images for gaze directions the training set contains thinly. 6
Quantitative evidence. Pre-training a tracker on two million synthetic image–gaze pairs and evaluating before any real-data fine-tuning, the smallest cohort in the study (0.5K subjects, 5M frames) gives an E50 gaze error of 0.64° for real-data-only training, 0.62° for the label-only control, and 0.59° for GazeDiT, with a paired-bootstrap p below 1.0 × 10⁻⁴ against real-only. The tail measure moves on the 1K-subject cohort from 2.74° to 2.62° (p = 1.7 × 10⁻⁷). After fine-tuning on real data, gaze error on difficult cases falls from 3.05° to 2.80° in the smallest cohort, which the authors present against a paired-bootstrap test over the same users. 6

Inspectability. The submission carries training and sampling configurations, the renderer description and the metric definitions in its appendices. It releases neither code nor the dataset, and the near-eye images come from an inward-facing VR-headset cohort that the paper describes as its own. 5
Caveat. The scope is single-frame near-eye generation with visible iris and pupil masks as the spatial representation, and the authors flag that extending the framework to richer geometry or temporally consistent gaze sequences is open. They also state that the label-plus-geometry principle will transfer to another domain only if that domain has an appropriate spatial representation, which makes the method's reach a per-domain research question. Reported gains are a few hundredths of a degree on the smallest cohort, and the largest effects require the tracker to be fine-tuned on real data afterwards. 6
Why it matters. Researchers generating labelled synthetic data for regression-style targets should read the conditioning interface here. The idea of routing a scalar label through the geometry in which it is physically expressed is the part that travels; eye tracking is the demonstration.
3. StrucPhysVideo: curating a corpus for physical events
Team and affiliation. The paper is submitted by the Awomo-WM Team and lists Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang and Kaicheng Yu as contributors, with correspondence routed through a
westlakedi.com address. The submission maps no author to an institution, so institutional attribution cannot be read off the paper. It is an arXiv v1 submission dated 16 September 2026. 7What changes. Video world models learn whatever their corpus contains, and web video is dominated by static product staging, presenter footage and post-production overlays. This paper treats that corpus as the problem and builds the model around it: a filtering and annotation pipeline that keeps clips where physical events are visible, plus structured captions that separate camera motion from object behaviour and name contact, deformation and state transitions explicitly. Two models follow from that data, a text-and-image-to-video model and an interactive image-and-action-to-video model driven by robot end-effector commands. 8
Technical insight. The trainable backbone is roughly 30 billion parameters across 48 blocks, with every feed-forward sublayer replaced by a sparse mixture-of-experts module that combines routed experts with one always-active shared expert. A frozen vision-language encoder jointly encodes prompt and reference image while a frozen VAE encodes the target video. The physical annotations use a two-level taxonomy of five top-level categories and 26 leaf phenomena. For the action-conditioned variant, a causal action encoder maps raw commands into latent action tokens, and a three-stage causalisation and distillation process converts the bidirectional teacher into a streaming student that runs four denoising steps per chunk with no classifier-free guidance. 8

Quantitative evidence. On Physics-IQ Verified the text-image-to-video model scores 45.5%, which the paper reports as the highest among compared models and 2.8 percentage points above Cosmos3-Super-Image2Video; the numbers for other models come from a benchmark snapshot dated 16 September 2026, and one entry marked with an asterisk is the authors' own reproduction. Caption ablations across backbones are reported as supporting the physics-focused supervision. The action-conditioned model is trained on AgiBot World-Beta plus more than 500 hours of general-domain video data and evaluated on a fixed subset of 5,000 clips from the AgiBotWorld-Alpha corpus against DreamDojo-AgiBot-14B and A2World. 8
Inspectability. The team released an inference and training code repository and a project page carrying generated samples in both the text-to-video and action-to-video settings. 910
Caveat. The reported headline rests on a single benchmark, Physics-IQ Verified, and the comparison mixes the authors' own reproduction with numbers taken from a benchmark snapshot. The evaluation of the action-conditioned model covers 5,000 clips from one robot corpus, so the embodied claim is narrower than the paper's framing suggests. This is a large training run: a 30-billion-parameter sparse mixture-of-experts backbone plus a distillation pipeline, which is a different order of resource from the four-step student that ships.
Why it matters. Researchers building video world models and embodied agents should read the data pipeline more closely than the headline score. The two-level physical taxonomy and the decision to caption camera motion separately from object behaviour are reusable on any corpus, whatever backbone is trained on the result.
4. DiT-Garment: 3D deformation learned in a 2D latent
Team and affiliation. Antoine Dumoulin, Laurence Boissieux and Stefanie Wuhrer at the Inria centre at the University Grenoble Alpes, Joao Regateiro at InterDigital Inc., and Pierre Hellier at Inria, University of Rennes, CNRS, IRISA-UMR 6074. The paper is an arXiv v1 submission dated 16 September 2026. 11
What changes. Animating a garment over a moving body has meant either a mesh graph network bound to a common template or a per-design model that must be retrained. DiT-Garment takes a garment design it has never seen, a motion it has never seen and a body shape it has never seen, and predicts dynamic deformation directly, producing a distribution of outcomes rather than one deterministic answer. 12
Technical insight. The deformation is learned as a 2D image problem in the garment's UV space. The template garment is a 3D triangle mesh aligned to a body in a standard pose; the model conditions on a 3D position map of that template rendered in UV space, so it implicitly learns how the 3D space around the body deforms without needing a shared template topology or graph convolutions. Body shape, a window of body motion and cloth material parameters enter as conditioning through cross-attention after a two-layer MLP, and the transformer uses linear attention with no positional encoding. A guidance strategy pushes the generated cloth out of the body. 12

Quantitative evidence. On a synthetic test set of 27 sequences where garment templates, motions and body shapes are all held out, DiT-Garment reports a mean vertex-to-vertex error of 5.20 cm against 6.95 cm for DPF, a Chamfer distance of 0.17 cm against 0.18 and 1.21, a crumple measure of 0.38 against 0.42 and 0.51, and 0.62% of cloth vertices inside the body against 0.64% for DPF. It stretches less than DPF, at 1.32 versus 0.22 in the stretch-error column where ContourCraft sits at 14.44, and runs at 0.4 frames per second against 0.05 for DPF and 1.7 for ContourCraft. The model trained only on simulations of automatically generated cloth designs then reproduces captured sequences from 4D-Dress and 4DHumanOutfit, and artist-made designs. 12
Inspectability. The authors host a project page with code and data for research use. 13
Caveat. This is a data-driven model bounded by the diversity and quality of its simulated training set, and the authors expect weaker generalisation for designs, bodies or motions well outside it. Physical constraints such as collision avoidance and energy conservation are enforced implicitly by the guidance rather than guaranteed, so intersections can still appear. Inference runs at 0.4 frames per second, and the model works non-autoregressively, which caps its use in interactive settings. 12
Why it matters. Graphics and 3D vision researchers working on avatars, virtual try-on and garment prototyping get a working example of a 2D generative prior doing 3D work: the UV-space formulation is what makes unseen designs generalise, and it is the transferable part of the paper.
5. VibeAvatar: two stages for two different qualities
Team and affiliation. Qilin Wang and Mingyu Li in the School of Computer Science and the School of Electronics Engineering and Computer Science at Peking University, with Hao Tang of the School of Computer Science at Peking University as corresponding author. The work is supported by the Fundamental Research Funds for the Central Universities. It is an arXiv v1 submission dated 16 September 2026, and the proceedings version appears at ACM Multimedia 2026. 1415
What changes. Talking-avatar systems are usually judged on two axes that pull against each other, lip synchronisation and perceived motion quality, and single-generator pipelines tend to trade one for the other. VibeAvatar splits them: articulation is shaped at the conditioning stage, and motion aesthetics are optimised afterwards, so post-training does not overwrite the articulation prior. 16
Technical insight. A Phonetic Kinematics Adapter converts recognition-oriented speech features into phonetic-kinematic conditions feeding a lightweight flow-based motion generator that works in a compact one-dimensional warp-based latent motion space. On top of that, an Aesthetic Motion Policy optimises a flow-consistent stochastic sampling policy with Group Relative Policy Optimization, using a timestep-truncated strategy, which is the step that aligns motion dynamics with human preference ratings without degrading synchronisation. 16
Quantitative evidence. Trained on HDTF and RAVDESS with unseen-speaker test splits, VibeAvatar reports HDTF FID 22.033, FVD 166.913, identity similarity 0.912, Sync-C 8.639, Sync-D 6.910, VQA 3.678 and ASE 2.152; on RAVDESS, FID 16.282, FVD 220.975, Sync-C 6.134 and ASE 2.577, the best or second-best column entry in each case against EchoMimic, FantasyTalking, FLOAT, Wan-S2V and Hallo3. For a 10-second 512-pixel video it reports 0.7B parameters, under 10 seconds of inference and roughly 3 GB of VRAM, against 3B/EchoMimic at about 8 minutes and 7 GB, 14B/Wan-S2V at about 25 minutes and 57 GB, and 10B/Hallo3 at about 50 minutes and 70 GB. A five-dimension user study covering lip sync, motion diversity, identity similarity, temporal smoothness and aesthetic perception is reported as ranking the method first in lip sync and aesthetic perception. 16

Inspectability. The arXiv version carries the method, tables and user study, and links no code or project page. The ACM Multimedia 2026 proceedings version is the peer-reviewed record. 14
Caveat. Both benchmarks are talking-face datasets with held-out speakers, and the efficiency figures are measured on one 10-second 512-pixel clip, so the resource comparison holds for that clip length rather than for arbitrary durations. The preprint states no limitations section of its own, which makes the scope of the aesthetic post-training hard to judge: the user study covers five perceptual dimensions, and anything outside them is unmeasured.
Why it matters. Digital-human and avatar researchers get a template for optimising a sampling policy with preference-based reinforcement learning while protecting an earlier conditioning stage, and the 0.7B-parameter operating point is close to the range where a talking-avatar model can run on a single consumer GPU.
What the five papers have in common
Each of these papers changes what its generator is coupled to during training, and each picks a different interface to do it. Contrastive Noise Alignment rewrites which noise vector a training image is paired with. GazeDiT routes a scalar gaze label through the pupil and iris geometry where that label physically lives. StrucPhysVideo rebuilds the training corpus so that clips showing physical events, with captions that name contact and deformation, are what the model sees. DiT-Garment conditions on a 3D position map rendered in UV space so a 2D transformer learns 3D deformation. VibeAvatar splits speech conditioning from an aesthetic post-training stage so the two stop competing. The connecting constraint is the coupling between the condition and the generated signal, and four of the five treat that coupling as something to be constructed rather than inherited.
A reading order follows from where the constraint sits:
- Start with Contrastive Noise Alignment if you train flow-matching models and want a change that costs nothing architecturally.
- Start with GazeDiT if you generate labelled synthetic data for a continuous target and need the label to survive generation.
- Start with StrucPhysVideo if you are building video world models or embodied agents and your bottleneck is the corpus.
- Start with DiT-Garment if you work on avatars, virtual try-on or 3D asset animation and need generalisation across unseen designs.
- Start with VibeAvatar if you work on talking heads or on preference-based post-training of a flow sampling policy.
References
- 1arXiv cs.CV new submissions
arxiv.org
- 2arXiv cs.LG new submissions
arxiv.org
- 3Contrastive Noise Alignment abstract
arxiv.org
- 4
- 5GazeDiT abstract
arxiv.org
- 6GazeDiT HTML paper
arxiv.org
- 7StrucPhysVideo abstract
arxiv.org
- 8StrucPhysVideo HTML paper
arxiv.org
- 9StrucPhysVideo code repository
github.com
- 10StrucPhysVideo project page
westlakedi-awomo.github.io
- 11DiT-Garment abstract
arxiv.org
- 12DiT-Garment HTML paper
arxiv.org
- 13DiT-Garment project page
dumoulina.github.io
- 14VibeAvatar abstract
arxiv.org
- 15
- 16VibeAvatar HTML paper
arxiv.org
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
