NTH

Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens

AuthorsAli Asaria, Tony Salomone, Deep Gandhi

July 7, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that a diffusion language model’s token order is not truly parallel or sequential, and that how you measure it can completely change the story.

Key results

686
probe suite size

Total prompts across six regimes

0.430-0.604
token-level τb range

Tie-aware Kendall τb for content-token commit order across non-JSON regimes

0.94-0.96
block-seq control τb

Pure position//16 block-sequential control for R1/R2/R3/R5

0.749
GSM8K AUROC

Negative entropy predicting correctness on math

What the paper found

This DeepMind/Google study audits the shipped open-weights checkpoint google/diffusiongemma-26B-A4B-it, a masked discrete-diffusion mixture-of-experts model built on Gemma 4, by instrumenting its EntropyBoundSampler.accept_canvas commit path rather than retraining anything. Across a 686-prompt probe suite spanning GSM8K, HumanEval, MBPP, factual recall, open-ended instructions, and constrained JSON, the authors find that DiffusionGemma is neither truly parallel nor purely sequential: content-token commit order shows a moderate left-to-right bias with tie-aware Kendall τb from 0.430 to 0.604, but it remains well below a block-sequential control near 0.94–0.96, and commit batches are large at roughly 13–26 tokens per accept-call with same-call tie rates up to 0.72. The headline claim is that granularity matters: block-τb rises smoothly as analysis bins grow from 4 to 64, with no special jump at 16, so 16 is just an analyst choice, not an architectural boundary. Behavior is also regime-dependent, with constrained JSON essentially order-independent at τb = −0.044, while commit confidence is meaningful only in some tasks: on GSM8K, negative entropy predicts correctness with AUROC 0.749, but on factual recall AUROC falls to 0.471. Despite a 48-step budget, generations finish in a small late burst at entropy 0.002–0.012 nats, well below the 0.1 bound, and overall accuracy stays comparable to the autoregressive sibling Gemma-4 26B-A4B-it on the scorable regimes.

Original abstract

Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured. We instrument DiffusionGemma 26B, a masked discrete-diffusion mixture-of-experts model built on Gemma 4, hooking its sampler's accept step to record which canvas positions commit, when, and at what confidence. Across a 686-prompt, six-regime probe suite we find that its decoding is neither parallel nor block-autoregressive: it follows a partial left-to-right commit bias whose apparent strength depends almost entirely on the granularity at which you look. Order is weak token by token and strengthens smoothly as the analysis is coarsened, so the model's "block size" turns out to be an artifact of the measuring ruler rather than the architecture. The model commits in large simultaneous batches, leaving much of the within-batch order genuinely undefined rather than merely unobserved. The behaviour is regime-dependent: structured JSON is committed in essentially arbitrary order, and a position's commit confidence tracks correctness on mathematical reasoning but carries no signal on factual recall. Commitment is aggressive, finishing in a short late burst well inside the step budget, while task accuracy matches the model's autoregressive Gemma-4 sibling. Beyond these findings, our central contribution is methodological: measuring decoding order honestly demands handling trailing-EOS padding, within-regime confounding, commit non-monotonicity, block-size sensitivity, and large commit-batch ties, each of which can otherwise manufacture a decoding-order result that is not really there.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis