Scaling Properties of Text Conditioning in Visual Generation
AuthorsZilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan
Resources
This work shows that better-structured language can make visual diffusion models more capable, and uses scaling laws to build a stronger prompting system.
Key results
Pearson correlation between Grounded Perplexity Gain and converged diffusion loss across the controlled caption sweep.
Power-law correlation between Effective Detailness and converged diffusion loss.
Score achieved by Ours (Qwen-Image), compared with 56.2 for the matched natural-language control.
Overall composition-and-reasoning score achieved by Ours (Qwen-Image).
Thinking-mode score using the 397B Qwen3.5 prompter, rising from 46.4% at 0.8B.
What the paper found
Researchers Zilong Chen and colleagues at ByteDance Seed study why simply making text prompts longer does not reliably improve text-to-image generation. Across 15 controlled caption configurations, they show that image-grounded information—not token count—predicts diffusion training quality. Their typed JSON structured prompt represents global scene context, object attributes, bounding boxes, depth, and relationships; information is measured with Grounded Perplexity Gain, or GPG, and Effective Detailness, or ED. On the BAGEL backbone, converged diffusion loss correlates with GPG at -0.984 and with ED at -0.971, while both measures rank caption configurations consistently. The system then combines structured supervision for higher diffusability with a Qwen3.5 prompter trained through supervised fine-tuning, cold-start distillation, and verifier-gated on-policy self-distillation, or OPSD. With Qwen-Image as the diffuser, the resulting ByteDance Seed system reaches 72.5 GenEval2 geometric mean and 85.2 on CoReBench, exceeding the matched natural-language Qwen-Image control, which scores 56.2 and 76.1. Zero-shot prompter scaling also transfers: thinking-mode GenEval++ rises from 46.4 percent with a 0.8B Qwen3.5 model to 86.8 percent with 397B. On broader comparisons, the system surpasses open-weight models and matches or exceeds closed systems such as OpenAI’s GPT-Image-1 and Google’s Nano Banana on most evaluations. An inference-time refine–render–judge loop provides additional gains, but performance saturates after a few rounds, indicating that training the prompt interface is more valuable than merely adding inference compute.
Original abstract
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.