Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
AuthorsShufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu, Aditya Grover, Jan Kautz, Pavlo Molchanov
Resources
This paper improves discrete diffusion image generation by letting the model revise earlier token choices and by training more effectively with large vocabularies, boosting both quality and efficiency.
Key results
decoder-only Nemotron-Labs-Diffusion-Image
Emu3.5 tokenizer codebook used in training
main benchmark score for text-to-image alignment
benchmark score on DPG
MJHQ-30k human preference score
faster than Emu3.5 in inference speed comparison
What the paper found
NVIDIA’s Nemotron-Labs-Diffusion-Image advances masked discrete diffusion for 1024px text-to-image synthesis by fixing two core failures in standard MDMs: irreversible token commitments and sparse supervision from very large codebooks. The model uses an 8B decoder-only transformer initialized from Nemotron-Labs-Diffusion and a 131,072-code Emu3.5 tokenizer, then adds token editing so already-unmasked tokens can be revised during inference when confidence exceeds 0.6. To stabilize training with large vocabularies, it introduces Grouped Cross-Entropy, which clusters discrete codes with K-means and adds hierarchical cross-entropy over 16,384 and 8,192 clusters, plus a fused custom operator that cuts GCE latency from 44.14 ms to 20.04 ms and max VRAM from 25.2 GB to 16.1 GB. Trained for 300K steps on 64 H100 GPUs over 16 days on 137M text-image pairs, the system reaches 0.90 on GenEval, 85.2 on DPG, and 9.61 HPSv3 on MJHQ-30k; with additional synthetic finetuning it rises to 86.9 DPG and 10.76 HPSv3. An ablation shows token editing makes 32 NFE sampling match no-edit 64 NFE quality, effectively halving forward calls, and the full model is reported as 42.4× faster than Emu3.5 at comparable or better GenEval quality.
Original abstract
We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard MDMs lack self-correcting capability because discrete tokens cannot be modified once they are unmasked. Second, although increasing the vocabulary size of discrete image tokenizers improves reconstruction fidelity, it introduces optimization difficulties for generative modeling as the per-token training signal becomes increasingly sparse. To address the first challenge, Nemotron-Labs-Diffusion-Image incorporates a token-editing mechanism that enables the model to dynamically revise already-unmasked tokens during inference, similar to how a sculptor iteratively refines their work. To tackle the second challenge, we propose a Grouped Cross-Entropy (GCE) objective that assigns positive learning signals to tokens neighboring the ground truth in embedding space, thereby alleviating signal sparsity. To further improve training efficiency, we implement a custom fused operator for GCE that significantly reduces VRAM usage in large-vocabulary settings. Experimental results demonstrate that these innovations substantially improve both training efficiency and image fidelity of masked discrete image generators, achieving a score of 0.90 on GenEval, 86.9 on DPG and 10.76 of HPSv3.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.