NTH

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

AuthorsShufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu, Aditya Grover, Jan Kautz, Pavlo Molchanov

July 2, 2026 2 min read
Watch on YouTube
The one-line take

This paper improves discrete diffusion image generation by letting the model revise earlier token choices and by training more effectively with large vocabularies, boosting both quality and efficiency.

Key results

8B
Model size

decoder-only Nemotron-Labs-Diffusion-Image

131,072
Tokenizer vocab

Emu3.5 tokenizer codebook used in training

0.90
GenEval overall

main benchmark score for text-to-image alignment

85.2
DPG

benchmark score on DPG

9.61
HPSv3

MJHQ-30k human preference score

42.4x
Latency reduction

faster than Emu3.5 in inference speed comparison

What the paper found

NVIDIA’s Nemotron-Labs-Diffusion-Image advances masked discrete diffusion for 1024px text-to-image synthesis by fixing two core failures in standard MDMs: irreversible token commitments and sparse supervision from very large codebooks. The model uses an 8B decoder-only transformer initialized from Nemotron-Labs-Diffusion and a 131,072-code Emu3.5 tokenizer, then adds token editing so already-unmasked tokens can be revised during inference when confidence exceeds 0.6. To stabilize training with large vocabularies, it introduces Grouped Cross-Entropy, which clusters discrete codes with K-means and adds hierarchical cross-entropy over 16,384 and 8,192 clusters, plus a fused custom operator that cuts GCE latency from 44.14 ms to 20.04 ms and max VRAM from 25.2 GB to 16.1 GB. Trained for 300K steps on 64 H100 GPUs over 16 days on 137M text-image pairs, the system reaches 0.90 on GenEval, 85.2 on DPG, and 9.61 HPSv3 on MJHQ-30k; with additional synthetic finetuning it rises to 86.9 DPG and 10.76 HPSv3. An ablation shows token editing makes 32 NFE sampling match no-edit 64 NFE quality, effectively halving forward calls, and the full model is reported as 42.4× faster than Emu3.5 at comparable or better GenEval quality.

Original abstract

We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard MDMs lack self-correcting capability because discrete tokens cannot be modified once they are unmasked. Second, although increasing the vocabulary size of discrete image tokenizers improves reconstruction fidelity, it introduces optimization difficulties for generative modeling as the per-token training signal becomes increasingly sparse. To address the first challenge, Nemotron-Labs-Diffusion-Image incorporates a token-editing mechanism that enables the model to dynamically revise already-unmasked tokens during inference, similar to how a sculptor iteratively refines their work. To tackle the second challenge, we propose a Grouped Cross-Entropy (GCE) objective that assigns positive learning signals to tokens neighboring the ground truth in embedding space, thereby alleviating signal sparsity. To further improve training efficiency, we implement a custom fused operator for GCE that significantly reduces VRAM usage in large-vocabulary settings. Experimental results demonstrate that these innovations substantially improve both training efficiency and image fidelity of masked discrete image generators, achieving a score of 0.90 on GenEval, 86.9 on DPG and 10.76 of HPSv3.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis