NTH

Abra: Scaling Diffusion Image Training

AuthorsKyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan

August 25, 2026 2 min read
Watch on YouTube
The one-line take

Abra shows that diffusion image models scale predictably, but reach their best performance with far more training data per parameter than language models.

Key results

2B
A BRA model range

The largest model in the A BRA scaling family; the family begins at 60M.

200
Compute-optimal image tokens per parameter

Recommended diffusion-training budget, compared with Chinchilla’s 20.

0.5%
Diffusion overtraining penalty

Loss penalty remains below this level at 2× the optimal token budget.

1B
Training dataset scale

DataComp-1B image-text dataset used for pretraining.

165
Optimal TPP at 256×256 resolution

Estimated compute-optimal image tokens per parameter.

247
Optimal TPP at 768×768 resolution

Estimated compute-optimal image tokens per parameter at higher resolution.

What the paper found

Abra Scaling Diffusion Image Training presents a compute-optimal scaling study for text-to-image diffusion using A BRA, a controlled family of latent flow-matching transformers spanning 60M to 2B parameters and trained with µP across 1e19 to 1e22 FLOPs. Using the DataComp-1B image-text dataset, a frozen Qwen3-4B text encoder, and a FLUX VAE, the study finds that diffusion training reaches compute optimality at approximately 200 image tokens per parameter, versus Chinchilla’s 20 tokens per parameter for language models—a 10x higher data requirement. Unlike LLMs such as Chinchilla-trained models, diffusion systems tolerate overtraining: at 2× the optimal token budget, the loss penalty is less than 0.5%, so a smaller model trained on more data is usually safer than a larger undertrained model. Loss, FID, KID, CLIPScore, CMMD, and ImageNet linear-probe accuracy follow predictable scaling laws, but their optimal data-to-parameter allocations differ; optimal CFG also decreases as model size and training compute increase. The paper reports the first scaling-collapse result for diffusion training, with differently sized models’ rescaled training curves converging to the noise floor within the first 15% of training. Resolution changes the rule: estimated optimal tokens per parameter rise from 165 at 256×256 pixels to 247 at 768×768, showing that higher-resolution generation is more data-intensive in token terms. The results provide a practical scaling prescription distinct from those used for OpenAI-style language models, DeepSeek systems, or image generators built from FLUX.

Original abstract

Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute ($10^{19}$ to $10^{22}$ FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately $200$ image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis