Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens
AuthorsAli Asaria, Tony Salomone, Deep Gandhi
Resources
This paper shows that a diffusion language model’s token order is not truly parallel or sequential, and that how you measure it can completely change the story.
Key results
Total prompts across six regimes
Tie-aware Kendall τb for content-token commit order across non-JSON regimes
Pure position//16 block-sequential control for R1/R2/R3/R5
Negative entropy predicting correctness on math
What the paper found
This DeepMind/Google study audits the shipped open-weights checkpoint google/diffusiongemma-26B-A4B-it, a masked discrete-diffusion mixture-of-experts model built on Gemma 4, by instrumenting its EntropyBoundSampler.accept_canvas commit path rather than retraining anything. Across a 686-prompt probe suite spanning GSM8K, HumanEval, MBPP, factual recall, open-ended instructions, and constrained JSON, the authors find that DiffusionGemma is neither truly parallel nor purely sequential: content-token commit order shows a moderate left-to-right bias with tie-aware Kendall τb from 0.430 to 0.604, but it remains well below a block-sequential control near 0.94–0.96, and commit batches are large at roughly 13–26 tokens per accept-call with same-call tie rates up to 0.72. The headline claim is that granularity matters: block-τb rises smoothly as analysis bins grow from 4 to 64, with no special jump at 16, so 16 is just an analyst choice, not an architectural boundary. Behavior is also regime-dependent, with constrained JSON essentially order-independent at τb = −0.044, while commit confidence is meaningful only in some tasks: on GSM8K, negative entropy predicts correctness with AUROC 0.749, but on factual recall AUROC falls to 0.471. Despite a 48-step budget, generations finish in a small late burst at entropy 0.002–0.012 nats, well below the 0.1 bound, and overall accuracy stays comparable to the autoregressive sibling Gemma-4 26B-A4B-it on the scorable regimes.
Original abstract
Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured. We instrument DiffusionGemma 26B, a masked discrete-diffusion mixture-of-experts model built on Gemma 4, hooking its sampler's accept step to record which canvas positions commit, when, and at what confidence. Across a 686-prompt, six-regime probe suite we find that its decoding is neither parallel nor block-autoregressive: it follows a partial left-to-right commit bias whose apparent strength depends almost entirely on the granularity at which you look. Order is weak token by token and strengthens smoothly as the analysis is coarsened, so the model's "block size" turns out to be an artifact of the measuring ruler rather than the architecture. The model commits in large simultaneous batches, leaving much of the within-batch order genuinely undefined rather than merely unobserved. The behaviour is regime-dependent: structured JSON is committed in essentially arbitrary order, and a position's commit confidence tracks correctness on mathematical reasoning but carries no signal on factual recall. Commitment is aggressive, finishing in a short late burst well inside the step budget, while task accuracy matches the model's autoregressive Gemma-4 sibling. Beyond these findings, our central contribution is methodological: measuring decoding order honestly demands handling trailing-EOS padding, within-regime confounding, commit non-monotonicity, block-size sensitivity, and large commit-batch ties, each of which can otherwise manufacture a decoding-order result that is not really there.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.