Register Tokens for Bounded-State Reasoning in Diffusion Language Models
AuthorsAlbert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala
Resources
The work gives diffusion language models a small set of learned memory tokens so they can preserve reasoning across multiple chunks without retaining the generated text.
Key results
Examples from OpenMathInstruct-2 and OpenCodeInstruct used for post-training.
Maximum register improvement over discrete-text carry on GSM8K.
Maximum register improvement over discrete-text carry on MBPP.
Bounded register carry versus full-context decoding in the reported cost comparison.
Reward-point improvement from registers over discrete text after chunked diffu-GRPO.
Held-out Pearson correlation when predicting the running total from register states.
What the paper found
This paper introduces register tokens, a fixed-size continuous memory mechanism for masked diffusion language models, allowing systems such as LLaDA-8B-Base and Dream-7B-Base to continue reasoning after each generated text chunk is erased. Four fixed-position embeddings are written from a completed chunk, carried into the next window, and rewritten; training uses prompt-masked continuation passes and a fully masked pass so the next chunk must depend on the registers rather than reconstructing the solution from visible context. Post-training on a 60K-example mixture of OpenMathInstruct-2 and OpenCodeInstruct evaluates 128-token math chunks and 64-token code chunks. Registers outperform four-token discrete-text carry by up to 8.5 points on GSM8K and 19.5 points on MBPP, with Dream reaching 40.9 on MBPP versus 21.4 for discrete text. Their advantage is strongest when outputs span multiple chunks: on LLaDA, register carry achieves 26.2 on HumanEval and 29.2 on MBPP, exceeding full-sequence SFT by 12.2 and 3.5 points. Although full-context decoding remains more accurate at a 1024-token horizon, bounded registers reduce decoding cost, reaching a 5.6× speedup in the reported comparison. The proposed chunked diffu-GRPO reinforcement-learning method further improves register state quality, raising LongArithmetic reward by 8.1 points over discrete carry; linear probes also decode the running total with Pearson R=0.84 and the next operation with 80.0% accuracy. The results position registers as compact, writable reasoning state rather than simple token memory.
Original abstract
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoning after that text is cleared, using only a fixed-size carried state. We implement this state as a small number of register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks. We post-train dLLMs to decode a chunk of text, clear it while preserving the register values, and continue decoding from the prompt and carried state. In our main comparisons on LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains of up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation, where correct programs usually span several chunks. Finally, registers can be further refined with reinforcement learning on long-horizon reasoning tasks.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.