Unlocking Lossless Speedups in LLMs via Discrete Diffusion
AuthorsSubham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
Resources
Uno uses discrete diffusion to make standard autoregressive LLMs generate multiple tokens in parallel, delivering up to three times faster inference without sacrificing quality.
Key results
Peak Uno speedup over the base autoregressive model
Uno tokens per second in the 1K/8K throughput test
Uno speedup at the largest batch size supported by the base AR model
Trainable diffusion LoRA parameters added to Qwen3-8B
Maximum end-to-end DAPO reinforcement-learning speedup
What the paper found
This paper introduces Uno, a diffusion-augmented LLM that preserves an autoregressive model’s exact output distribution while drafting multiple tokens in parallel. Its architecture keeps the original AR weights responsible for quality and adds lightweight rank-128 LoRA diffusion adapters trained with Diffusion Distillation, combining Discrete Consistency Distillation with a total-variation objective. The Ψ-Spec sampler then verifies diffusion drafts using standard rejection sampling, requiring no separate draft model unlike EAGLE-3 or DFlash and avoiding the quality loss associated with diffusion models such as Google’s DiffusionGemma and NVIDIA’s Nemotron-Labs-Diffusion. An 8B Uno model reaches 5255 tokens per second in system throughput, versus 3577 for its base AR model, while retaining a 1.5× speedup at the largest supported batch size and up to 3× speedup in favorable settings. On Qwen3-8B, UnoQwen adds only 0.35B parameters, exceeds 5700 tokens per second at high concurrency, and achieves up to 2.5× per-request acceleration. Quality remains unchanged because the AR verifier defines the final distribution: Uno scores 68.4 on SWE-bench Verified, 68.0 on AA-LCR, and 90.1 on τ2 Telecom, outperforming the 26B DiffusionGemma and proprietary Mercury 2 across agentic, coding, and long-context benchmarks. The same adapters also accelerate DAPO reinforcement-learning rollouts, delivering up to 40% end-to-end training speedup.
Original abstract
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $Ψ$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.