NTH

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

AuthorsDengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi

August 25, 2026 2 min read
Watch on YouTube
The one-line take

This study shows how to efficiently train pixel-space image diffusion models by first learning in latent space, achieving competitive quality with substantially faster inference.

Key results

20B
Pre-training corpus

Image-text pairs used for the controlled large-scale comparison

0.20
Z-Image-Turbo pixel latency

Seconds per 1024-by-1024 image on an NVIDIA H800

4.75
Z-Image-Turbo speedup

End-to-end speedup over latent Z-Image-Turbo

0.7698
Z-Image-Turbo GenEval

GenEval score for the distilled pixel-space model

86.85
Z-Image-Turbo DPG

DPG score for the distilled pixel-space model

10.09M
DiP decoder parameters

Decoder parameter count in the quality-efficiency comparison

What the paper found

This study from Alibaba examines how to make pixel-space text-to-image diffusion competitive with latent-space systems such as Z-Image and FLUX2-klein. Using more than 20B image–text pairs, it finds that direct RGB pre-training converges substantially more slowly because the model must learn both global structure and high-frequency statistics without a compact VAE representation. The proposed solution is latent-to-pixel adaptation: pre-train in latent space, initialize the pixel model from those weights, mix source-model-generated images with real images, switch to x-prediction, use the lightweight convolutional DiP decoder, calibrate the noise scale to gamma equals 2, and progressively enlarge patches from ps16 to ps32. Larger patches reduce visual tokens fourfold, while Decoupled-DMD and DMDR step distillation reduce sampling to 4 function evaluations. On an NVIDIA H800 generating 1024-by-1024 images, the resulting Z-Image-Turbo pixel model reaches a GenEval score of 0.7698 and DPG score of 86.85 at 0.20 seconds per image, a 4.75 times speedup over latent Z-Image-Turbo. The DiP decoder itself uses only 10.09M parameters and 834.03 GFLOPs, offering a practical quality-efficiency balance. The recipe also transfers to FLUX2-klein, indicating that pixel-space generation can avoid VAE decoding latency without sacrificing overall benchmark performance.

Original abstract

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis