NTH

Scaling Properties of Text Conditioning in Visual Generation

AuthorsZilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan

August 7, 2026 2 min read
Watch on YouTube
The one-line take

This work shows that better-structured language can make visual diffusion models more capable, and uses scaling laws to build a stronger prompting system.

Key results

-0.984
GPG-to-diffusion-loss correlation

Pearson correlation between Grounded Perplexity Gain and converged diffusion loss across the controlled caption sweep.

-0.971
ED-to-diffusion-loss correlation

Power-law correlation between Effective Detailness and converged diffusion loss.

72.5
GenEval2 geometric mean

Score achieved by Ours (Qwen-Image), compared with 56.2 for the matched natural-language control.

85.2
CoReBench

Overall composition-and-reasoning score achieved by Ours (Qwen-Image).

86.8%
GenEval++ at largest prompter scale

Thinking-mode score using the 397B Qwen3.5 prompter, rising from 46.4% at 0.8B.

What the paper found

Researchers Zilong Chen and colleagues at ByteDance Seed study why simply making text prompts longer does not reliably improve text-to-image generation. Across 15 controlled caption configurations, they show that image-grounded information—not token count—predicts diffusion training quality. Their typed JSON structured prompt represents global scene context, object attributes, bounding boxes, depth, and relationships; information is measured with Grounded Perplexity Gain, or GPG, and Effective Detailness, or ED. On the BAGEL backbone, converged diffusion loss correlates with GPG at -0.984 and with ED at -0.971, while both measures rank caption configurations consistently. The system then combines structured supervision for higher diffusability with a Qwen3.5 prompter trained through supervised fine-tuning, cold-start distillation, and verifier-gated on-policy self-distillation, or OPSD. With Qwen-Image as the diffuser, the resulting ByteDance Seed system reaches 72.5 GenEval2 geometric mean and 85.2 on CoReBench, exceeding the matched natural-language Qwen-Image control, which scores 56.2 and 76.1. Zero-shot prompter scaling also transfers: thinking-mode GenEval++ rises from 46.4 percent with a 0.8B Qwen3.5 model to 86.8 percent with 397B. On broader comparisons, the system surpasses open-weight models and matches or exceeds closed systems such as OpenAI’s GPT-Image-1 and Google’s Nano Banana on most evaluations. An inference-time refine–render–judge loop provides additional gains, but performance saturates after a few rounds, indicating that training the prompt interface is more valuable than merely adding inference compute.

Original abstract

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis