Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
AuthorsJiayu Yang, Chao Chen, Shengen Wu, Yinhong Liu, Yuxuan Fan, Lujundong Li, Songning Lai, Chengwei Qin, Zhijiang Guo
Resources
This paper introduces SWITCH, a way to turn latent reasoning on and off with special tokens so large language models can be trained with reinforcement learning while keeping their hidden reasoning more interpretable.
Key results
Switch after Switch-GRPO on the main benchmark
Absolute improvement over the strongest Coconut-style baseline at matched scale
Switch after Switch-GRPO on the arithmetic benchmark
Increase from curriculum-only checkpoint after Switch-GRPO on MATH-500
Switch usage before RL in the curriculum-only checkpoint
Last-layer linear probe accuracy for predicting that the next token is <swi>
What the paper found
This paper from HKUST(GZ), the University of Cambridge, and NTU introduces Switch, a switchable latent-reasoning framework for Qwen3-8B that makes hidden-state recurrence compatible with on-policy reinforcement learning by adding explicit <swi> and </swi> boundary tokens around latent blocks. The key technical move is that GRPO is applied only at discrete text positions, while latent positions are replayed through deterministic hidden-state injection, so the policy ratio remains well-defined during training. Switch is trained in three phases: supervised tagging of high-entropy reasoning spans, a curriculum that progressively replaces text with <latent> steps, and a Switch-GRPO objective with correctness, format, latent-usage, and optional brevity rewards. On MATH-500, the full system reaches 79.3% accuracy, beating the strongest Coconut-style hidden-state recurrence baseline by 25.7 points at the same Qwen3-8B scale, and on GSM8K it reaches 89.2%. Relative to the curriculum-only checkpoint, Switch-GRPO raises latent-conditional accuracy by 12.6 points while cutting switch usage from 81% to 58%. Mechanistic analysis shows that <swi> is a sharply localized learned control token, probe accuracy for predicting the switch decision reaches 91.9% at the last layer, and causal interventions confirm that the injected latent state is functionally necessary: zeroing it drops diagnostic accuracy from 100% to 33.3%, whereas a same-norm random vector is far less harmful. The authors conclude that hidden-state recurrence is both RL-trainable and directly interpretable.
Original abstract
Latent chain-of-thought compresses reasoning by replacing visible reasoning traces with continuous hidden-state recurrence, but existing formulations are difficult to optimize with standard on-policy reinforcement learning (RL) and hard to interpret causally. Our key insight is that a single pair of explicit boundary tokens can address both issues at once: discrete entry and exit anchors make the latent block compatible with standard on-policy RL, and the same anchors offer a natural foothold for mechanistic analysis. Motivated by this, we propose SWITCH, a switchable latent reasoning framework. The model emits <swi> to enter latent mode and </swi> to exit. Because the boundaries are ordinary discrete tokens, the GRPO policy ratio is well-defined at every decision point. The same anchors also expose the latent steps to direct probing and causal intervention. We train the model with a visible-to-latent curriculum and a Switch-GRPO objective that propagates gradients through recurrent latent computation. SWITCH consistently outperforms prior hidden-state-recurrence latent reasoning approaches at similar scale. Mechanistic analysis through the boundary tokens further reveals three findings: (i) <swi> is a sharply localised, learned switching policy rather than a stylistic artefact; (ii) the latent step it opens performs problem-specific, causally important computation rather than acting as an inert placeholder; and (iii) that computation is concentrated at a single hidden-state transition on entry. Together, these results show that hidden-state-recurrence latent reasoning is both RL-trainable and open to direct mechanistic analysis, including of how on-policy RL itself improves the model from the inside.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.