DiffusionGemma Technical Report
AuthorsDiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
Resources
DiffusionGemma turns a large language model into a fast parallel text generator, reaching roughly 1,500 tokens per second while preserving much of the model's reasoning and multimodal capability.
Key results
DiffusionGemma’s mixture-of-experts model size.
Tokens denoised in parallel per block.
Average adaptive denoising steps across evaluations.
Average diffusion decoding efficiency across seven benchmarks.
Approximate output tokens per second on one NVIDIA H100.
DiffusionGemma text-diffusion score in thinking mode.
What the paper found
Google DeepMind’s DiffusionGemma is an experimental open-weight language model that replaces strictly sequential autoregressive decoding with discrete multinomial diffusion. Fine-tuned from the Gemma 4 26B A4B mixture-of-experts model, it uses 25.2B total parameters and 3.85B activated parameters, denoising 256-token canvases with bidirectional attention before appending them to a causal KV cache. Its two-stage training pipeline combines supervised fine-tuning with sampler distillation and reinforcement learning, using less than 10% of the original Gemma 4 training-token budget. An entropy-bounded sampler, temperature annealing, self-conditioning, and adaptive stopping reduce the typical denoising trajectory to 12 steps, while the model averages 19.74 tokens per forward pass and reaches 1500 output tokens per second on a single NVIDIA H100 at batch size one. On the evaluation suite, it scores 73.2 on GPQA-Diamond and 69.1 on LiveCodeBench-v6, trading some capability against Gemma 4’s autoregressive mode for a 7.1x throughput improvement over standard Gemma 4 decoding and a 4.8x improvement over its multi-token-prediction variant. DiffusionGemma retains Gemma 4’s thinking mode, multimodal inputs, long-context support, and autoregressive operation, while its bidirectional refinement enables self-correction and rapid convergence on structured outputs. The main limitations are occasional repetition, concise reasoning traces, lower peak quality than the autoregressive baseline, and weaker throughput at high batch sizes.
Original abstract
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.