NTH

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

AuthorsChuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie

September 11, 2026 2 min read
Watch on YouTube
The one-line take

LLaDA-Image offers an open, diffusion-based recipe for building highly capable image generators and distills them into a model that can generate and edit images in just a few steps.

Key results

6B
DiT model size

Diffusion Transformer trained from scratch

220M
Generation-training samples

Total samples processed across the image-generation pipeline

98%
Real-image share

Proportion of the 220M generation-training samples consisting of real images

53.53
Qwen-Image-Bench English

Overall English-track score

53.38
Qwen-Image-Bench Chinese

Overall Chinese-track score

4
Turbo benchmark sampling

Sampling steps used for LLaDA-Image Turbo benchmark evaluation

What the paper found

LLaDA-Image presents a fully open recipe for a unified text-to-image and instruction-guided editing model. Its core is a 6B-parameter Diffusion Transformer trained from scratch, paired with a frozen LLaDA2.0-Mini vision-language module using SigLIP-VQ features. Instead of beginning with expensive image-text pairs, the system learns its visual prior through masked image-only pre-training and mid-training, then adds bilingual text alignment, high-resolution refinement, and joint generation-editing supervision. The 220M-sample pipeline is 98% real images, while parameter-free RMSNorm and the Muon optimizer stabilize long training. Editing combines semantic reference tokens with clean FLUX.2 VAE latents to preserve identity, structure, texture, and unchanged regions. TwinFlow distills the full model into LLaDA-Image Turbo for 2–4 sampling steps; benchmark evaluation uses four steps. On Qwen-Image-Bench, LLaDA-Image scores 53.53 in English and 53.38 in Chinese, establishing open-source state of the art over models such as Z-Image Turbo and Qwen-Image. Comparisons also include proprietary systems including OpenAI’s GPT-Image 2 and Google DeepMind’s Nano-Banana 2.0. The release includes model weights, training and inference code, and detailed data-processing and optimization recipes, although editing perceptual quality and long, multi-region text rendering remain weaker than the best specialized systems.

Original abstract

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis