NTH

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

AuthorsFuning Fu, Tenghui Wang, Junyong Cen, Qichao Zhu, Guanyu Zhou

June 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that in latent diffusion models, predicting the clean latent directly can work better than predicting velocity, because the choice of training target changes the geometry of learning in compressed latent space.

Key results

130M
Parameters

trainable JLT Base-scale model size

250K
Training steps

shared training schedule for JLT and matched velocity baseline

2.56
FID-50K (JLT-B/1 vs DiT-B/1)

clean-latent prediction in matched latent ablation

6.56
FID-50K (DiT-B/1)

velocity prediction in matched latent ablation

14.81
FID-50K (JLT-B/2 vs DiT-B/2)

clean-latent prediction with stronger token aggregation

28.71
FID-50K (DiT-B/2)

velocity prediction with stronger token aggregation

What the paper found

JLT: Clean Latent Prediction in Latent Diffusion Transformers isolates a single question inside latent diffusion: if generation already happens in a compressed FLUX.2 VAE space, does the direct prediction target still matter? Using a controlled 130M-parameter Base-scale Transformer with the same 12-block, 768-hidden-dimension backbone, the same 250K-step training schedule, and the same ImageNet 256×256 evaluation protocol, the paper compares clean-latent prediction x against matched velocity prediction v. The result is a large, consistent gap: JLT-B/1 improves FID-50K from 6.56 to 2.56 and Inception Score from 132.12 to 220.74, while JLT-B/2 improves FID-50K from 28.71 to 14.81. With classifier-free guidance, the final JLT-B/1 model reaches FID-50K 2.50 and IS 232.51, versus 14.00 FID-50K without guidance. The authors’ local Gaussian analysis explains why the algebraically equivalent targets are not equivalent for finite models: velocity regression adds an isotropic covariance floor, inflates low-variance latent directions, and yields larger conditional residuals than clean prediction. In short, the paper argues that target parameterization in latent diffusion is a geometric modeling choice, not a notation change, and that clean-latent regression is substantially easier to optimize in a fixed latent space.

Original abstract

Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predicting an ambient noised quantity. We ask whether this principle remains useful after images are mapped into a learned latent space, where compression has already removed much of the raw pixel variability. We introduce JLT, a 130M latent diffusion Transformer over frozen FLUX.2 VAE codes, and compare clean-latent prediction with a matched velocity-prediction DiT under the same representation, backbone, and training settings. Although the three variables x, epsilon, and v are linearly convertible for a fixed corruption time, a local Gaussian analysis shows that velocity regression inherits an isotropic target-covariance floor and amplifies low-variance latent directions, while clean prediction damps them. On ImageNet 256 x 256, JLT-B/1 obtains FID-50K 2.50 with classifier-free guidance, with a large matched-target gap over velocity prediction. These results suggest that prediction targets in latent diffusion are representation-dependent geometric choices, rather than interchangeable algebraic parameterizations.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis