NTH

Register Tokens for Bounded-State Reasoning in Diffusion Language Models

AuthorsAlbert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala

September 18, 2026 2 min read
Watch on YouTube
The one-line take

The work gives diffusion language models a small set of learned memory tokens so they can preserve reasoning across multiple chunks without retaining the generated text.

Key results

60K
Training mixture

Examples from OpenMathInstruct-2 and OpenCodeInstruct used for post-training.

8.5
GSM8K gain

Maximum register improvement over discrete-text carry on GSM8K.

19.5
MBPP gain

Maximum register improvement over discrete-text carry on MBPP.

5.6×
Decoding speedup

Bounded register carry versus full-context decoding in the reported cost comparison.

8.1
LongArithmetic RL gain

Reward-point improvement from registers over discrete text after chunked diffu-GRPO.

0.84
Running-total probe

Held-out Pearson correlation when predicting the running total from register states.

What the paper found

This paper introduces register tokens, a fixed-size continuous memory mechanism for masked diffusion language models, allowing systems such as LLaDA-8B-Base and Dream-7B-Base to continue reasoning after each generated text chunk is erased. Four fixed-position embeddings are written from a completed chunk, carried into the next window, and rewritten; training uses prompt-masked continuation passes and a fully masked pass so the next chunk must depend on the registers rather than reconstructing the solution from visible context. Post-training on a 60K-example mixture of OpenMathInstruct-2 and OpenCodeInstruct evaluates 128-token math chunks and 64-token code chunks. Registers outperform four-token discrete-text carry by up to 8.5 points on GSM8K and 19.5 points on MBPP, with Dream reaching 40.9 on MBPP versus 21.4 for discrete text. Their advantage is strongest when outputs span multiple chunks: on LLaDA, register carry achieves 26.2 on HumanEval and 29.2 on MBPP, exceeding full-sequence SFT by 12.2 and 3.5 points. Although full-context decoding remains more accurate at a 1024-token horizon, bounded registers reduce decoding cost, reaching a 5.6× speedup in the reported comparison. The proposed chunked diffu-GRPO reinforcement-learning method further improves register state quality, raising LongArithmetic reward by 8.1 points over discrete carry; linear probes also decode the running total with Pearson R=0.84 and the next operation with 80.0% accuracy. The results position registers as compact, writable reasoning state rather than simple token memory.

Original abstract

Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoning after that text is cleared, using only a fixed-size carried state. We implement this state as a small number of register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks. We post-train dLLMs to decode a chunk of text, clear it while preserving the register values, and continue decoding from the prompt and carried state. In our main comparisons on LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains of up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation, where correct programs usually span several chunks. Finally, registers can be further refined with reinforcement learning on long-horizon reasoning tasks.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis