NTH

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

AuthorsSubham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu

September 11, 2026 2 min read
Watch on YouTube
The one-line take

Uno uses discrete diffusion to make standard autoregressive LLMs generate multiple tokens in parallel, delivering up to three times faster inference without sacrificing quality.

Key results

3
Maximum speedup

Peak Uno speedup over the base autoregressive model

5255
System throughput

Uno tokens per second in the 1K/8K throughput test

1.5
Largest-batch speedup

Uno speedup at the largest batch size supported by the base AR model

0.35B
UnoQwen added parameters

Trainable diffusion LoRA parameters added to Qwen3-8B

40%
RL training speedup

Maximum end-to-end DAPO reinforcement-learning speedup

What the paper found

This paper introduces Uno, a diffusion-augmented LLM that preserves an autoregressive model’s exact output distribution while drafting multiple tokens in parallel. Its architecture keeps the original AR weights responsible for quality and adds lightweight rank-128 LoRA diffusion adapters trained with Diffusion Distillation, combining Discrete Consistency Distillation with a total-variation objective. The Ψ-Spec sampler then verifies diffusion drafts using standard rejection sampling, requiring no separate draft model unlike EAGLE-3 or DFlash and avoiding the quality loss associated with diffusion models such as Google’s DiffusionGemma and NVIDIA’s Nemotron-Labs-Diffusion. An 8B Uno model reaches 5255 tokens per second in system throughput, versus 3577 for its base AR model, while retaining a 1.5× speedup at the largest supported batch size and up to 3× speedup in favorable settings. On Qwen3-8B, UnoQwen adds only 0.35B parameters, exceeds 5700 tokens per second at high concurrency, and achieves up to 2.5× per-request acceleration. Quality remains unchanged because the AR verifier defines the final distribution: Uno scores 68.4 on SWE-bench Verified, 68.0 on AA-LCR, and 90.1 on τ2 Telecom, outperforming the 26B DiffusionGemma and proprietary Mercury 2 across agentic, coding, and long-context benchmarks. The same adapters also accelerate DAPO reinforcement-learning rollouts, delivering up to 40% end-to-end training speedup.

Original abstract

Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $Ψ$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis