High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
AuthorsDongyang Liu, Ruoyi Du, David Liu, Dengyang Jiang, Liangchen Li, Qilong Wu, Zhen Li, Steven C. H. Hoi, Hongsheng Li, Peng Gao
Resources
This paper shows how to squeeze high-quality images out of a diffusion model in just two denoising steps by aligning the student to its teacher, separating step-specific parameters, and training end-to-end with regularization.
Key results
Overall score of the final 2-step model
Overall score of the final 2-step model
Overall score of the final 2-step model
Final 2-step model score
Final 2-step model score
Score when the step-1 loss is removed
What the paper found
Alibaba Group’s Z-Image Team and The Chinese University of Hong Kong present Z-Image Turbo++, a 2-step image generator distilled from the 8-step Z-Image Turbo teacher and built on the 6B-parameter Z-Image S3-DiT foundation. The paper targets the hard limit of few-step diffusion by combining Distribution-Aligned Adversarial Learning, Step-Decoupled Parameterization, and End-to-End Training with Iterative Regularization. The key novelty is to use 8-step teacher-generated images as GAN “real” samples instead of external photos, which the authors report is more stable and yields cleaner outputs; they also split the two denoising steps into independently updated weights, effectively doubling step-specific capacity. In evaluation on OneIGBench, GenEval, DPGBench, and LongTextBench, the final 2-step model reaches 52.50 on OneIGBench, 75.70 on GenEval, 85.86 on DPGBench, and 91.62/89.88 on LongText-CN/LongText-EN, substantially narrowing the gap to the 8-step teacher while preserving sharp texture and stronger text rendering than prior 2-step baselines such as TwinFlow and DMD2. The ablations show that removing the step-1 loss drops OneIGBench to 50.16 and LongText-CN to 84.49, confirming that iterative regularization is critical for keeping the first step meaningful. Training uses 16 H100 GPUs, 20,000 iterations, and about 80 hours, with a memory-efficient backpropagation trick that passes the step-2 gradient into step 1 via an inherited loss.
Original abstract
Few-step diffusion distillation has become increasingly mature for 4-8-step generation, yet pushing further to 2 steps remains challenging. In this work, we introduce Z-Image Turbo++, a high-quality 2-step image generation model distilled from the 8-step Z-Image Turbo teacher. Our method addresses the central bottlenecks of increased task difficulty and limited model capacity in 2-step generation through three simple but effective design choices tailored to this regime. First, we propose Distribution-Aligned Adversarial Learning, which uses teacher-generated images rather than external real images as real samples for GAN training, providing a more attainable and informative adversarial target. Second, we adopt Step-Decoupled Parameterization, assigning independent model parameters to the two denoising steps to better match their distinct capacity demands. Third, we perform End-to-End Training with Iterative Regularization, allowing the first step to receive gradients from final image quality while preserving a meaningful intermediate generation through an explicit step-1 loss. Together, these designs substantially narrow the quality gap between 2-step and 8-step generation in both qualitative and quantitative evaluations, highlighting the potential of carefully tailored distillation strategies for improving the quality-efficiency trade-off in few-step generation.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.