JLT: Clean-Latent Prediction in Latent Diffusion Transformers
AuthorsFuning Fu, Tenghui Wang, Junyong Cen, Qichao Zhu, Guanyu Zhou
Resources
This paper shows that in latent diffusion models, predicting the clean latent directly can work better than predicting velocity, because the choice of training target changes the geometry of learning in compressed latent space.
Key results
trainable JLT Base-scale model size
shared training schedule for JLT and matched velocity baseline
clean-latent prediction in matched latent ablation
velocity prediction in matched latent ablation
clean-latent prediction with stronger token aggregation
velocity prediction with stronger token aggregation
What the paper found
JLT: Clean Latent Prediction in Latent Diffusion Transformers isolates a single question inside latent diffusion: if generation already happens in a compressed FLUX.2 VAE space, does the direct prediction target still matter? Using a controlled 130M-parameter Base-scale Transformer with the same 12-block, 768-hidden-dimension backbone, the same 250K-step training schedule, and the same ImageNet 256×256 evaluation protocol, the paper compares clean-latent prediction x against matched velocity prediction v. The result is a large, consistent gap: JLT-B/1 improves FID-50K from 6.56 to 2.56 and Inception Score from 132.12 to 220.74, while JLT-B/2 improves FID-50K from 28.71 to 14.81. With classifier-free guidance, the final JLT-B/1 model reaches FID-50K 2.50 and IS 232.51, versus 14.00 FID-50K without guidance. The authors’ local Gaussian analysis explains why the algebraically equivalent targets are not equivalent for finite models: velocity regression adds an isotropic covariance floor, inflates low-variance latent directions, and yields larger conditional residuals than clean prediction. In short, the paper argues that target parameterization in latent diffusion is a geometric modeling choice, not a notation change, and that clean-latent regression is substantially easier to optimize in a fixed latent space.
Original abstract
Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predicting an ambient noised quantity. We ask whether this principle remains useful after images are mapped into a learned latent space, where compression has already removed much of the raw pixel variability. We introduce JLT, a 130M latent diffusion Transformer over frozen FLUX.2 VAE codes, and compare clean-latent prediction with a matched velocity-prediction DiT under the same representation, backbone, and training settings. Although the three variables x, epsilon, and v are linearly convertible for a fixed corruption time, a local Gaussian analysis shows that velocity regression inherits an isotropic target-covariance floor and amplifies low-variance latent directions, while clean prediction damps them. On ImageNet 256 x 256, JLT-B/1 obtains FID-50K 2.50 with classifier-free guidance, with a large matched-target gap over velocity prediction. These results suggest that prediction targets in latent diffusion are representation-dependent geometric choices, rather than interchangeable algebraic parameterizations.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.