Five diffusion papers from the September 23 batch: malign overfitting, point clouds at any resolution, and a production editing system

Five diffusion papers from the September 23 batch: malign overfitting, point clouds at any resolution, and a production editing system

A ranked scan of five diffusion-model preprints from the September 23, 2026 arXiv batch: a production e-commerce editing system, a theory of why diffusion overfitting is malign, a rule for splitting flow-matching schedules, SE(3)-equivariant point-cloud generation at any resolution, and perceptual distance read off a denoising trajectory.

arXiv's newest announcement is the batch for Wednesday, 23 September 2026: 108 new cs.CV submissions and 118 new cs.LG submissions, covering the submission cycle that closed with arXiv's 22 September 2026 cutoff. 12 A listing page groups everything announced together, and a few of those 226 entries carry v1 dates from earlier in the week, so each candidate was checked individually against the arXiv API. All five papers below carry verified v1 timestamps on 22 September 2026 (UTC), the latest at 13:30.
Citation counts do not exist yet at this age, so they cannot rank anything. Three routes were tried on this run's papers: an OpenAlex lookup by arXiv DOI returned HTTP 404 for the selected identifiers, Semantic Scholar's graph API returned HTTP 429, and GitHub's repository API — the one new route — returned 14 stars for the KwaiMind repository and 1 star for the FoMo repository, which is an attention signal on released code rather than on the papers themselves. The ranking below rests on method novelty, author-reported evidence, disclosed affiliation and public inspectability. Every number in this digest is author-reported unless a sentence says otherwise.

The five at a glance

RankPaperCentral moveStrongest reported signalInspectability
1KwaiMindAn agent-run data engine plus staged post-training that tunes an editing model for e-commerceStrongest overall among the open-source editors it compares against on ImgEdit, GEdit, both REDEdit splits and its own Ecom-Bench; CTR-guided optimization lifts the share of outputs whose predicted CTR beats the original product image from 12.16% to 37.41% 3Code released
2Double Descent and Malign Overfitting in Diffusion ModelsSeparating what the number of noise draws per sample and the number of samples each controlInterpolation peak at p ~ nm, but the test loss starts rising at p ~ n; optimally regularized large models beat unregularized ones of any size 4No code in the preprint
3Gaussian Flow-Matching SchedulesSplitting a flow-matching schedule into a variance path and a factorizationClosed-form factorizations that minimize or equalize time-averaged regression variance along any fixed path; a necessary drift bound for exact N-step Euler sampling 5No code in the preprint
4EMERGEThe first fully SE(3)-equivariant graph-based diffusion backbone for point clouds1-NNA EMD 50.71 on ShapeNet cars against a theoretical optimum of 50, and 2,600–7,000 epochs to converge where LION reports 32,000 6No code in the preprint
5FoMoReading a perceptual distance off the timestep where two denoising trajectories forkBest average SROCC on all four reference-based IQA benchmarks, trained on 480k generated pairs and no human labels 7Code released

1. KwaiMind: a data engine and three reward models for commerce editing

Team and affiliation. The KwaiMind Team at Kuaishou Group. The submission is a technical report, arXiv v1 dated 22 September 2026. 3
What changes. Commercial image editing asks for three things a general editor does not optimize for: the product has to stay the same product, any text has to render correctly, and the result has to be appealing enough to sell. KwaiMind builds the data pipeline, the training stages and the reward models around those three requirements, and reports results on general editing benchmarks and on a new e-commerce benchmark with a click-through-rate ranking. 8
Technical insight. The model follows the architectural design of Qwen-Image and is a multimodal diffusion transformer (MMDiT) with three concatenated input streams. The Data Agent maintains roughly 1.8M high-quality training pairs, drawn down from 26.2M raw pairs contributed by ScaleEdit-12M (12.0M), X2Edit (3.7M), AnyEdit (2.5M) and the team's own KwaiData (8.0M). Training runs in stages: continued pre-training on the subset above 720p with instruction augmentation; supervised fine-tuning on a task-balanced corpus of 115.0K general and 119.9K e-commerce pairs; direct preference optimization in a mixed offline and online regime; then online reinforcement learning on the forward diffusion process, driven by a vision-language judge for general edits and by three dedicated reward models that target what a general reward misses — a click-through-rate model for commercial appeal, a coarse-to-fine reward for text rendering, and a consistency reward for product and model identity. 8
A four-phase data agent system: pre-filter routing samples through voter and judge aggregation, a generation agent handling modification and supplementation, a caption agent producing training labels, and a post-filter with a simulated human evaluator, with a human-in-the-loop band beneath
Figure 4 of the report. The bounded feedback loops are the mechanism rather than the plumbing: Pre-Filter sends recoverable rejects back into the Generation Agent, and that recovery is why the maintained corpus ends at 1.8M pairs instead of the 26.2M the sources offered. 8
Quantitative evidence. Ecom-Bench covers 11 commercial editing tasks with 100 test samples each, scored on four task-specific dimensions on a 1–5 rubric and combined as a geometric mean, plus a CTR-based ranking. KwaiMind takes the strongest overall scores among the evaluated open-source editors on ImgEdit, GEdit, both the English and Chinese splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among the compared systems. CTR-guided optimization raises the share of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images produced about a 2.44% relative increase in actual CTR. 8
Inspectability. The training and inference code is released at KwaiMmu/KwaiMind. The repository was created on 25 August 2026 and last pushed on 23 September 2026; no model weights are announced. 9
Caveat. The strongest-overall claim is among open-source editors: closed-source systems appear in the same figures as hatched bars and are not beaten. The benchmark, the CTR model and the reward models are all introduced by the same team that reports the gains. The report's own next steps are closing the visual-quality gap with proprietary systems and improving robustness on compositional and multi-reference edits, so those are the two places to test the system first. 8
Why it matters. Researchers working on controllable editing and industry teams shipping one get a complete recipe, and the transferable piece is the reward set. Three narrow rewards — predicted CTR, text rendering, identity consistency — replace the single generic preference signal that a general-purpose DPO or RLHF stage would use, which is a pattern worth copying wherever a domain has its own measurable objective. The code is public.

2. Double Descent: the peak is not where the damage starts

Team and affiliation. Raphaël Urfin and Giulio Biroli at the Laboratoire de Physique de l'École normale supérieure (ENS, Université PSL, CNRS, Sorbonne Université, Université Paris Cité); Tony Bonnaire at the Institut d'Astrophysique Spatiale, Université Paris-Saclay and CNRS; Marc Mézard at the Department of Computing Sciences, Bocconi University. arXiv v1 dated 22 September 2026. 4
What changes. Training a diffusion model is training a regression — a quadratic denoising score-matching loss — and regression is the setting where overparameterization acquired its reputation for being benign. Diffusion models nevertheless overfit and memorize their training sets. This paper resolves the contradiction by working out where the double-descent curve's landmarks actually fall once the training procedure's own variable is put back into the analysis. 10
Technical insight. The paper carries two instruments. The first is closed-form learning curves for a random-feature neural network in the proportional limit, where dimension, sample count and parameter count all go to infinity together at fixed ratios; the analysis extends earlier work to finite m, the number of noise realizations drawn per training sample. The second is experiments: DDPM-style U-Nets trained on CelebA grayscale images downsampled to 32×32, with n from 512 to 8192, each image assigned m fixed noise realizations, network widths W from 2 (about 29K parameters) to 192 (about 240M), for 2M Adam steps. Separating m from n is what changes the picture. 10
Three panels of double-descent curves: loss against parameters per dimension for the random-feature model, loss against parameter count for U-Nets trained on CelebA, and the plateau test loss against diffusion time
Figure 2 of the paper. The middle panel is the one to read against your own runs: the U-Nets, trained on 2,048 CelebA images with 8 noise draws each and evaluated at t ≈ 0.01, peak and then keep rising well past the parameter count where the random-feature curve has already turned down. 10
Quantitative evidence. The interpolation peak does occur, but at p ~ nm rather than the p ~ n of standard regression. The rise of the test loss sets in far earlier, at p ~ n, and does not depend on m. A bias–variance decomposition locates the mechanism: the bias of the score estimator begins to grow at p ~ n, keeps growing past the interpolation peak, and saturates at a large value, while the variance decays as it does in regression. Because diffusion models are trained with m ≫ 1, the peak is pushed to very large model sizes, which places practical training runs on the rising branch before it, where the overfitting is already active. Regularization restores the usual benefit of scale: ridge-regularized random-feature models and early-stopped U-Nets outperform unregularized models of any size. 10
Caveat. The empirical side is one dataset, a moderate number of samples, and models far smaller than the state of the art. The random-feature model and the U-Net do not behave identically — they differ in how the interpolation peak height depends on n, and the U-Net shows a bias peak the random-feature model lacks — so the authors leave the role of feature learning to future work. The whole analysis is on 32×32 grayscale images. 10
Inspectability. No code, data or checkpoint link appears in the preprint. 10
Why it matters. The practical reading is not "do not scale up". It is that memory and memorization are governed by the bias of the score estimator, which grows monotonically with width from p ~ n onward, and that this growth is what the ridge penalty or early stopping is there to hold down. If you are deciding where to spend capacity on a diffusion model, this paper is the argument for spending some of it on regularization instead, and it says which half of the error to watch while you do.

3. Gaussian Flow-Matching Schedules: the path and the parameterization are two decisions

Team and affiliation. Arsène Claustre and Kimia Nadjahi at CNRS, ENS Paris; Hugo Negrel at CNRS, ENS Paris and the Machine Learning Lab, Capital Fund Management; Claire Boyer at the Laboratoire de Mathématiques d'Orsay, CNRS, Université Paris-Saclay, and the Institut Universitaire de France; Eric Vanden-Eijnden at CNRS, ENS Paris, Capital Fund Management, and the Courant Institute of Mathematical Sciences, New York University. arXiv v1 dated 22 September 2026. 5
What changes. Flow matching fixes where a path starts and where it ends and leaves the schedule — how the two endpoints are mixed at intermediate times — entirely to the designer. That single choice is asked to do two jobs. At sampling it determines the probability path, and therefore how accurately a finite number of Euler steps can follow it. At training it determines the regression target and the variance of that target. This paper shows the two jobs can be assigned separately. 11
Technical insight. For centered, commuting Gaussians and a direction-dependent schedule, the schedule splits into a variance path, which fully determines the intermediate laws and the probability flow, and a factorization, which leaves that flow untouched while controlling the irreducible regression variance. On the sampling side the authors analyze finite-step Euler accuracy and derive a necessary drift bound for exact N-step sampling, which connects the geodesic path and the logarithmic path as two ends of the same condition. On the training side they derive, for any fixed path, closed-form factorizations that either minimize the time-averaged regression variance or hold it constant along the path. 11
Two panels: optimal eigenvalue against time for several regularization strengths, and optimal alpha and beta schedules against time for three objectives
Figure 1 of the paper. The right panel is the result a practitioner can use: kinetic energy, the integrated squared Jacobian and the variance criterion each select a different schedule, so "the best schedule" is only meaningful once you say which criterion you are optimizing. 11
Quantitative evidence. There is no benchmark here, and the paper does not claim one. The evidence is the derivation plus the numerical illustration in Figure 1: the optimal eigenvalue r_t at several regularization strengths, spanning kinetic energy at λ = 0 to a Lipschitz criterion as λ → ∞, and the α_t and β_t schedules selected by the three objectives. 11
Caveat. The exact results hold for centered, commuting Gaussians under a direction-dependent schedule. The authors state plainly that beyond that setting the decomposition should be read as a covariance-informed design principle rather than an exact characterization, and that the necessary drift bound is a condition on the path rather than a recipe for building one. Nothing in the paper is measured on an image model, so the transfer to a trained generative model on real data is untested here. 11
Inspectability. No code link appears in the preprint. 11
Why it matters. If you have been tuning a noise schedule and watching training variance move in the opposite direction from sampling quality, this is the paper that says why the two moved together, and separates them. As a theory paper it is fast to read: the decomposition in Section 2 carries most of the value, and Figure 1 tells you which of the three criteria your problem is closest to.

4. EMERGE: point clouds that ignore their own resolution

Team and affiliation. Ilias Mitsouras, Nikolaos Chaidos, Giorgos Stamou and Athanasios Voulodimos, all at the National Technical University of Athens, Athens, Greece. arXiv v1 dated 22 September 2026. 6
What changes. Point-cloud generation has been built on Transformers and VAEs and then on diffusion placed on top of them, and those backbones either ignore the continuous spatial symmetries of 3D data or approximate them with discrete rotation groups. EMERGE puts an equivariant graph network at the centre of the diffusion backbone, so that rotating or translating the input rotates or translates the output exactly rather than approximately. The consequence a user feels is that a model trained at one point count generates at other point counts without retraining or fine-tuning. 12
Technical insight. The denoiser is a stack of blocks, each holding a multi-scale EGNN that aggregates local features across several neighbourhood scales, followed by canonical voxel pooling and a global invariant feature attention module. Each point's message carries a 13-dimensional invariant feature vector — local curvature, linearity, planarity — computed from the eigenvalues and eigenvectors of its neighbourhood covariance, which adds geometric context without breaking equivariance. The network predicts the clean point cloud rather than the noise, with Min-SNR weighting at γ = 5. Resolution changes at inference are handled by a parameter-free distribution alignment: a cloud's point density changes with its point count, the noise level that matches a given density shifts with it, and the alignment negates that shift analytically instead of retraining. 12
The EMERGE architecture: a noisy point cloud entering a stack of multi-scale equivariant graph blocks with voxel pooling and global attention, producing a denoised cloud, with an inset showing canonical-frame projection
Figure 1 of the paper. The pooling inset on the right is where the canonical frame is established; that step is what has to work before the equivariance claim means anything, and it is also where the paper's computational cost comes from. 12
Quantitative evidence. Evaluation uses 1-NNA, where 50% is the theoretical optimum, with both Chamfer distance and Earth Mover's distance, on the ShapeNet airplane, chair and car categories at 100 DDIM steps. Under the standard split EMERGE reaches 55.93 CD and 53.33 EMD on airplanes, 54.31 and 52.42 on chairs, and 53.55 and 50.71 on cars, ahead of every baseline listed, including LION at 67.41 and 61.23 for airplanes, 53.70 and 52.34 for chairs, and 53.41 and 51.14 for cars. Under the LION split it leads in five of the six metrics. Training converges in 3,100 epochs for airplanes, 2,600 for chairs and 7,000 for cars, against LION's 32,000 total epochs (24,000 for its latent diffusion model plus 8,000 for its VAE). Zero-shot super-resolution is demonstrated up to 40,000 points without any adaptation. 12
Caveat. Wall-clock time moves against the epoch count. The k-nearest-neighbour graph is rebuilt at every layer, and the invariant features need analytical eigenvector computations for both the canonical frame and the feature vectors, neither of which has the hardware-level GPU support that standard attention enjoys. The authors state that this makes the time per training step and per inference pass higher than for standard diffusion baselines, so fewer epochs do not automatically mean faster training. The preprint carries no code or project page. 12
Why it matters. For anyone generating 3D assets or doing reconstruction, the resolution-agnostic inference is the part that changes a pipeline: train on 2,048-point clouds, generate 40,000-point ones, with no fine-tuning. The equivariance argument is also a clean test case for a broader question — whether a strong geometric inductive bias still pays for itself when per-step cost is higher.

5. FoMo: the denoising trajectory as a perceptual ruler

Team and affiliation. Jaihyun Lew and Minjun Park in the Interdisciplinary Program in AI, Mingi Jung and Wooseok Song in the Department of Electrical and Computer Engineering, and Sungroh Yoon in both plus the AIIS, ASRI, INMC and ISRC institutes, all at Seoul National University. arXiv v1 dated 22 September 2026. 7
What changes. Reference-based image quality metrics learn from human labels, and the standard label is a two-alternative forced choice between two distorted images. That is cheap to collect and reliable per pair, but it only ever encodes relative comparisons, never a global ordering. FoMo replaces the annotation with the diffusion model's own generation process: forward-diffuse a reference image to a chosen timestep, denoise it back, and the timestep at which two generations stop sharing a trajectory measures how far apart the results are. A fork early in the process leaves only coarse structure in common and means the images are perceptually far apart; a fork late means they differ in fine detail. 13
Technical insight. The forking step becomes a pointwise distance label, which turns the training objective from pairwise classification into a ranking problem, supervised here with a RankNet-style binary cross-entropy. The paper generates 480k training pairs, drawing forking steps uniformly from s ~ U[0, 49] and using SD 1.5, SD-XL and SD3 so that the reference pool covers real and synthetic domains. A first-order Taylor expansion shows why the pointwise label carries more: the ranking signal grows linearly with the gap between forking steps, where a binary vote encodes only its sign. 13
The FoMo framework: a reference image and its forking trajectories producing labelled samples, beside the binary comparison matrix of 2AFC supervision and the ranked comparison matrix of the proposed objective
Figure 2 of the paper. The left half is the whole idea in one picture — the vertical gap between two branching trajectories is the distance label. The right half is why the label is worth more than a vote: the upper matrix is the pairwise outcome 2AFC gives you, the lower one is the ordering FoMo synthesises. 13
Quantitative evidence. A human study with 30 participants, 1,982 responses and 190 reference images supports the link between forking moment and perceived distance. Table 1 reports mean SROCC over five random seeds across four reference-based IQA benchmarks and seven backbones. On PIPAL, training on FoMo's generated labels reaches an average of 0.683 against 0.484 for the best human-annotated dataset, PieAPP; on TID2013, 0.719 against 0.700 for KADID-10K; on CSIQ, 0.874 against 0.794 for KADID-10K; on LIVE, 0.926 against 0.870 for KADID-10K. FoMo leads the average column on all four benchmarks. 13
Inspectability. Code is released at JHLew/FoMo. 14
Caveat. The distorted images are made by diffusion models, so the authors state that a metric trained this way can fail on image domains outside that distribution, artificial distortions in particular, and their proposed remedy is to train a diffusion model on the new domain and regenerate the labels there. That makes the pipeline's coverage a property of whichever generative model is used to build it. The comparison in Table 1 is between label sources — FoMo's generated labels against four human-annotated datasets — rather than between finished metrics, and the human study verifies the trajectory-to-perception link on the study's own reference images. 13
Why it matters. Anyone who trains a restoration model, a super-resolver or a perceptual loss needs a distance that tracks human judgement, and human labelling is the usual bottleneck. This paper turns a diffusion model already on the machine into the annotator, and the released code includes the generation pipeline rather than only the trained metric. The transferable idea is the fork itself: the timestep at which two samples stop sharing a trajectory is a quantity you can read off any diffusion model without a human in the loop.

What the five have in common

Three of the five take a design choice that is normally made as one decision and split it in two. Gaussian Flow-Matching Schedules separates the path, which sets sampling quality, from the factorization, which sets the variance of the training target. Double Descent separates where overfitting begins, at p ~ n, from where the interpolation peak sits, at p ~ nm. EMERGE separates the resolution a model is trained at from the resolution it generates at, which is what the equivariant backbone and the density alignment buy together. The other two add a signal instead of separating one: KwaiMind adds three measurable domain rewards on top of a general preference stage, and FoMo adds a label source in place of human annotation. Each split or addition is the paper's actual contribution, and each one is the thing to check first when deciding whether the result transfers.
The batch itself put three of the five in cs.CV and two in cs.LG, and the split runs through subject matter rather than method: the three vision papers are about producing or judging an image or a shape, the two machine-learning papers are about the training and sampling process that produces anything.
A reading order follows from which of those decisions is currently open in your own work:
  • Start with KwaiMind if you are adapting a diffusion editor to a commercial domain and need a pattern for reward design rather than another architecture.
  • Start with Double Descent if you are choosing model size and regularization for a diffusion model and want the argument for why the two have to be chosen together.
  • Start with Gaussian Flow-Matching Schedules if you are designing or debugging a noise schedule and have been treating sampling accuracy and training variance as one knob.
  • Start with EMERGE if you generate point clouds or meshes and want inference resolution decoupled from training resolution.
  • Start with FoMo if you need perceptual distances or quality labels at a scale human annotation cannot reach.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel