Abra: Scaling Diffusion Image Training
AuthorsKyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan
Resources
Abra shows that diffusion image models scale predictably, but reach their best performance with far more training data per parameter than language models.
Key results
The largest model in the A BRA scaling family; the family begins at 60M.
Recommended diffusion-training budget, compared with Chinchilla’s 20.
Loss penalty remains below this level at 2× the optimal token budget.
DataComp-1B image-text dataset used for pretraining.
Estimated compute-optimal image tokens per parameter.
Estimated compute-optimal image tokens per parameter at higher resolution.
What the paper found
Abra Scaling Diffusion Image Training presents a compute-optimal scaling study for text-to-image diffusion using A BRA, a controlled family of latent flow-matching transformers spanning 60M to 2B parameters and trained with µP across 1e19 to 1e22 FLOPs. Using the DataComp-1B image-text dataset, a frozen Qwen3-4B text encoder, and a FLUX VAE, the study finds that diffusion training reaches compute optimality at approximately 200 image tokens per parameter, versus Chinchilla’s 20 tokens per parameter for language models—a 10x higher data requirement. Unlike LLMs such as Chinchilla-trained models, diffusion systems tolerate overtraining: at 2× the optimal token budget, the loss penalty is less than 0.5%, so a smaller model trained on more data is usually safer than a larger undertrained model. Loss, FID, KID, CLIPScore, CMMD, and ImageNet linear-probe accuracy follow predictable scaling laws, but their optimal data-to-parameter allocations differ; optimal CFG also decreases as model size and training compute increase. The paper reports the first scaling-collapse result for diffusion training, with differently sized models’ rescaled training curves converging to the noise floor within the first 15% of training. Resolution changes the rule: estimated optimal tokens per parameter rise from 165 at 256×256 pixels to 247 at 768×768, showing that higher-resolution generation is more data-intensive in token terms. The results provide a practical scaling prescription distinct from those used for OpenAI-style language models, DeepSeek systems, or image generators built from FLUX.
Original abstract
Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute ($10^{19}$ to $10^{22}$ FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately $200$ image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.