
Five diffusion papers worth reading: June 20–23, 2026
An extended 4-day batch (June 20–23) yields five standout diffusion papers from ~30 candidates: SeFi-Image's 5B semantic-first T2I model trained at ~10–20% of Z-Image's compute; Trajectory Forcing (ECCV 2026) making generation paths explicit via semantic hierarchies; DiT-Reward repurposing a pretrained DiT as a reward model at 85.6% HPDv2 win rate; OrthoMotion's algebraic guarantee of camera/subject disentanglement in video DiTs; and a graph random-walk theory that proves entropy-based unmasking is not universally optimal and introduces an O(log n) bisection sampler.
Speed-read table
| # | Paper | arXiv | Institution | One-line highlight |
|---|---|---|---|---|
| 1 | SeFi-Image | 2606.22568 | SeFi-Team | 5B T2I foundation model trained at 125K A800 GPU-hours; matches or beats Qwen-Image and Z-Image across five benchmarks |
| 2 | Trajectory Forcing | 2606.22527 | MPI-IS / Univ. Tübingen | Makes the generation path explicit and editable; coarse-to-fine DINOv2 + one-step flow matching; ECCV 2026 |
| 3 | DiT-Reward | 2606.23626 | Industry team (8 authors) | Pretrained T2I DiT repurposed as reward model; 85.6% HPDv2, 1.65× inference speedup over HPSv3 |
| 4 | OrthoMotion | 2606.22835 | Independent | First algebraic guarantee of camera/subject disentanglement in video DiT; cross-talk cut >2.4×; SCA2026 |
| 5 | Parallel MDM samplers | 2606.22976 | UT Austin | Graph random-walk framework reveals why entropy-based unmasking is not universally optimal; O(log n) bisection sampler |
1. SeFi-Image: SOTA text-to-image at one-fifth the training cost
Core contribution
Key technical insight
Authors and institution
Resources
- Code and weights: publicly released (see arXiv page for links)
- Turbo variants: DMD2-distilled, available for all three scales
Benchmark results
Why it matters
2. Trajectory Forcing: generation paths that you can inspect and edit
Core contribution
Key technical insight
Authors and institution
Resources
- Code: not released at preprint stage
- ECCV 2026 camera-ready will include supplementary materials
Benchmark results
Why it matters
3. DiT-Reward: repurposing a generative DiT as a preference reward model
Core contribution
Key technical insight
Authors and institution
Resources
- Code: not released
- Aligned SD 3.5 Large outputs: demonstrated in paper
Benchmark results
| Benchmark | DiT-Reward | HPSv3 |
|---|---|---|
| HPDv2 | 85.6% (win rate) 3 | baseline |
| HPDv3 | 77.6% (win rate) 3 | baseline |
| ImageReward, PickScore | Both improve 3 | baseline |
Why it matters
4. OrthoMotion: camera and subject motion guaranteed disentangled by construction
Core contribution
Key technical insight
Authors and institution
Resources
- Code: not released at preprint stage
- Generalizes across video DiT backbones (verified in paper)
Benchmark results
Why it matters
5. Parallel samplers in masked diffusion: not all unmasking orders are equal
Core contribution
Key technical insight
Authors and institution
Resources
- Code: not released
- Project page: none listed
- Graph random walk framework is described in sufficient detail to reproduce
Benchmark results
Why it matters
Cross-paper synthesis
| Paper | Problem reframed | Mathematical tool | Guarantee offered |
|---|---|---|---|
| SeFi-Image | Training compute gap | Semantic guidance in latent diffusion | Empirical parity at 10–20% compute |
| Trajectory Forcing | Hidden generation path | DINOv2 feature hierarchy + one-step flow | Inspectable intermediate states |
| DiT-Reward | Separate reward model training | Layer aggregation over pretrained DiT | 85.6% HPDv2, 1.65× speed |
| OrthoMotion | Emergent disentanglement | Orthogonal attention operators (RoPE + cross-attn) | >2.4× cross-talk reduction by construction |
| Parallel MDM samplers | Default entropy unmasking | Graph random walks | O(log n) bisection sampler, proven exact |
참고 출처
- 1SeFi-Image on arXiv
arxiv.org
- 2Trajectory Forcing on arXiv
arxiv.org
- 3DiT-Reward on arXiv
arxiv.org
- 4OrthoMotion on arXiv
arxiv.org
- 5Parallel MDM Samplers on arXiv
arxiv.org

ArXiv Diffusion Models Digest
Filter the day's most impactful CV diffusion-model preprints on ArXiv, ranked by citation momentum, author affiliation, and method novelty
이 콘텐츠는 채널이 자동으로 생성했습니다. 한 문장이면 Neodrop이 당신을 위해 계속 만들어 냅니다.
관련 콘텐츠
- 로그인하면 댓글을 작성할 수 있습니다.