NTH

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

AuthorsYonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov

July 8, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a language model that can switch between autoregressive, diffusion, and self-speculative decoding to trade off speed and quality, achieving major throughput gains over current open-source models.

Key results

70.28
25B-token base average after full pipeline

Average accuracy after block-wise attention, global loss averaging, DP-rank varying masking ratios, two-stage training, and AR loss

0.3
best diffusion weight alpha

Joint AR-diffusion coefficient giving the best balance

63.61
Nemotron-Labs-Diffusion-8B instruct average accuracy

Average over 10 tasks in instruct evaluation

2.57
diffusion TPF

Tokens per forward for Nemotron-Labs-Diffusion-8B diffusion mode

5.99
linear self-speculation TPF

Tokens per forward for LoRA-tuned linear self-speculation

4
SPEED-Bench throughput gain

Throughput improvement on SPEED-Bench with SGLang on NVIDIA GB200

What the paper found

NVIDIA’s Nemotron-Labs-Diffusion introduces a tri-mode language model that unifies autoregressive decoding, block-wise diffusion decoding, and self-speculation in one architecture, trained with a joint AR-plus-diffusion objective. The paper’s core claim is that AR and diffusion are complementary: adding the diffusion loss preserves or slightly improves AR accuracy, while the AR loss materially improves diffusion planning and stability. On 25B-token continuous pretraining, the ablation pipeline rises from 54.23% average accuracy with block-wise attention to 70.28% after adding global loss averaging, DP-rank varying masking ratios, two-stage training, and AR loss; the best diffusion weight is 𝛼=0.3. The released Nemotron-Labs-Diffusion-8B reaches 63.61% average accuracy in instruct mode, and its diffusion mode decodes 2.57 tokens per forward pass while its LoRA-tuned linear self-speculation reaches 5.99 tokens per forward. Against Qwen3-8B, the model’s AR mode is +0.86% better on average, its diffusion mode is +0.43% better, and on SPEED-Bench with SGLang on an NVIDIA GB200 GPU it delivers 4× higher throughput. A speed-of-light analysis shows a 7.60× diffusion upper bound on SPEED-Bench, 76.5% above linear self-speculation’s realized token-per-forward rate, indicating substantial remaining sampler headroom. The family scales to 3B, 8B, and 14B, extends to vision-language models, and uses a 20M-trajectory sampler trained with a 4-layer, 384-hidden-dimension Transformer.

Original abstract

We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis