LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
AuthorsChuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie
Resources
LLaDA-Image offers an open, diffusion-based recipe for building highly capable image generators and distills them into a model that can generate and edit images in just a few steps.
Key results
Diffusion Transformer trained from scratch
Total samples processed across the image-generation pipeline
Proportion of the 220M generation-training samples consisting of real images
Overall English-track score
Overall Chinese-track score
Sampling steps used for LLaDA-Image Turbo benchmark evaluation
What the paper found
LLaDA-Image presents a fully open recipe for a unified text-to-image and instruction-guided editing model. Its core is a 6B-parameter Diffusion Transformer trained from scratch, paired with a frozen LLaDA2.0-Mini vision-language module using SigLIP-VQ features. Instead of beginning with expensive image-text pairs, the system learns its visual prior through masked image-only pre-training and mid-training, then adds bilingual text alignment, high-resolution refinement, and joint generation-editing supervision. The 220M-sample pipeline is 98% real images, while parameter-free RMSNorm and the Muon optimizer stabilize long training. Editing combines semantic reference tokens with clean FLUX.2 VAE latents to preserve identity, structure, texture, and unchanged regions. TwinFlow distills the full model into LLaDA-Image Turbo for 2–4 sampling steps; benchmark evaluation uses four steps. On Qwen-Image-Bench, LLaDA-Image scores 53.53 in English and 53.38 in Chinese, establishing open-source state of the art over models such as Z-Image Turbo and Qwen-Image. Comparisons also include proprietary systems including OpenAI’s GPT-Image 2 and Google DeepMind’s Nano-Banana 2.0. The release includes model weights, training and inference code, and detailed data-processing and optimization recipes, although editing perceptual quality and long, multi-region text rendering remain weaker than the best specialized systems.
Original abstract
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.