An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
AuthorsDengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi
Resources
This study shows how to efficiently train pixel-space image diffusion models by first learning in latent space, achieving competitive quality with substantially faster inference.
Key results
Image-text pairs used for the controlled large-scale comparison
Seconds per 1024-by-1024 image on an NVIDIA H800
End-to-end speedup over latent Z-Image-Turbo
GenEval score for the distilled pixel-space model
DPG score for the distilled pixel-space model
Decoder parameter count in the quality-efficiency comparison
What the paper found
This study from Alibaba examines how to make pixel-space text-to-image diffusion competitive with latent-space systems such as Z-Image and FLUX2-klein. Using more than 20B image–text pairs, it finds that direct RGB pre-training converges substantially more slowly because the model must learn both global structure and high-frequency statistics without a compact VAE representation. The proposed solution is latent-to-pixel adaptation: pre-train in latent space, initialize the pixel model from those weights, mix source-model-generated images with real images, switch to x-prediction, use the lightweight convolutional DiP decoder, calibrate the noise scale to gamma equals 2, and progressively enlarge patches from ps16 to ps32. Larger patches reduce visual tokens fourfold, while Decoupled-DMD and DMDR step distillation reduce sampling to 4 function evaluations. On an NVIDIA H800 generating 1024-by-1024 images, the resulting Z-Image-Turbo pixel model reaches a GenEval score of 0.7698 and DPG score of 86.85 at 0.20 seconds per image, a 4.75 times speedup over latent Z-Image-Turbo. The DiP decoder itself uses only 10.09M parameters and 834.03 GFLOPs, offering a practical quality-efficiency balance. The recipe also transfers to FLUX2-klein, indicating that pixel-space generation can avoid VAE decoding latency without sacrificing overall benchmark performance.
Original abstract
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.