NTH

DiffusionGemma Technical Report

AuthorsDiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor

August 7, 2026 2 min read
Watch on YouTube
The one-line take

DiffusionGemma turns a large language model into a fast parallel text generator, reaching roughly 1,500 tokens per second while preserving much of the model's reasoning and multimodal capability.

Key results

25.2B
Total parameters

DiffusionGemma’s mixture-of-experts model size.

256
Canvas length

Tokens denoised in parallel per block.

12
Effective denoising steps

Average adaptive denoising steps across evaluations.

19.74
Tokens per forward

Average diffusion decoding efficiency across seven benchmarks.

1500
H100 output speed

Approximate output tokens per second on one NVIDIA H100.

73.2
GPQA-Diamond score

DiffusionGemma text-diffusion score in thinking mode.

What the paper found

Google DeepMind’s DiffusionGemma is an experimental open-weight language model that replaces strictly sequential autoregressive decoding with discrete multinomial diffusion. Fine-tuned from the Gemma 4 26B A4B mixture-of-experts model, it uses 25.2B total parameters and 3.85B activated parameters, denoising 256-token canvases with bidirectional attention before appending them to a causal KV cache. Its two-stage training pipeline combines supervised fine-tuning with sampler distillation and reinforcement learning, using less than 10% of the original Gemma 4 training-token budget. An entropy-bounded sampler, temperature annealing, self-conditioning, and adaptive stopping reduce the typical denoising trajectory to 12 steps, while the model averages 19.74 tokens per forward pass and reaches 1500 output tokens per second on a single NVIDIA H100 at batch size one. On the evaluation suite, it scores 73.2 on GPQA-Diamond and 69.1 on LiveCodeBench-v6, trading some capability against Gemma 4’s autoregressive mode for a 7.1x throughput improvement over standard Gemma 4 decoding and a 4.8x improvement over its multi-token-prediction variant. DiffusionGemma retains Gemma 4’s thinking mode, multimodal inputs, long-context support, and autoregressive operation, while its bidirectional refinement enables self-correction and rapid convergence on structured outputs. The main limitations are occasional repetition, concise reasoning traces, lower peak quality than the autoregressive baseline, and weaker throughput at high batch sizes.

Original abstract

We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis