NTH

Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning

AuthorsJiayu Yang, Chao Chen, Shengen Wu, Yinhong Liu, Yuxuan Fan, Lujundong Li, Songning Lai, Chengwei Qin, Zhijiang Guo

June 28, 2026 3 min read
Watch on YouTube
The one-line take

This paper introduces SWITCH, a way to turn latent reasoning on and off with special tokens so large language models can be trained with reinforcement learning while keeping their hidden reasoning more interpretable.

Key results

79.3
MATH-500 accuracy

Switch after Switch-GRPO on the main benchmark

25.7
MATH-500 gain over Coconut baseline

Absolute improvement over the strongest Coconut-style baseline at matched scale

89.2
GSM8K accuracy

Switch after Switch-GRPO on the arithmetic benchmark

12.6
Latent-conditional accuracy gain

Increase from curriculum-only checkpoint after Switch-GRPO on MATH-500

81
Switch rate reduction

Switch usage before RL in the curriculum-only checkpoint

91.9
Probe accuracy

Last-layer linear probe accuracy for predicting that the next token is <swi>

What the paper found

This paper from HKUST(GZ), the University of Cambridge, and NTU introduces Switch, a switchable latent-reasoning framework for Qwen3-8B that makes hidden-state recurrence compatible with on-policy reinforcement learning by adding explicit <swi> and </swi> boundary tokens around latent blocks. The key technical move is that GRPO is applied only at discrete text positions, while latent positions are replayed through deterministic hidden-state injection, so the policy ratio remains well-defined during training. Switch is trained in three phases: supervised tagging of high-entropy reasoning spans, a curriculum that progressively replaces text with <latent> steps, and a Switch-GRPO objective with correctness, format, latent-usage, and optional brevity rewards. On MATH-500, the full system reaches 79.3% accuracy, beating the strongest Coconut-style hidden-state recurrence baseline by 25.7 points at the same Qwen3-8B scale, and on GSM8K it reaches 89.2%. Relative to the curriculum-only checkpoint, Switch-GRPO raises latent-conditional accuracy by 12.6 points while cutting switch usage from 81% to 58%. Mechanistic analysis shows that <swi> is a sharply localized learned control token, probe accuracy for predicting the switch decision reaches 91.9% at the last layer, and causal interventions confirm that the injected latent state is functionally necessary: zeroing it drops diagnostic accuracy from 100% to 33.3%, whereas a same-norm random vector is far less harmful. The authors conclude that hidden-state recurrence is both RL-trainable and directly interpretable.

Original abstract

Latent chain-of-thought compresses reasoning by replacing visible reasoning traces with continuous hidden-state recurrence, but existing formulations are difficult to optimize with standard on-policy reinforcement learning (RL) and hard to interpret causally. Our key insight is that a single pair of explicit boundary tokens can address both issues at once: discrete entry and exit anchors make the latent block compatible with standard on-policy RL, and the same anchors offer a natural foothold for mechanistic analysis. Motivated by this, we propose SWITCH, a switchable latent reasoning framework. The model emits <swi> to enter latent mode and </swi> to exit. Because the boundaries are ordinary discrete tokens, the GRPO policy ratio is well-defined at every decision point. The same anchors also expose the latent steps to direct probing and causal intervention. We train the model with a visible-to-latent curriculum and a Switch-GRPO objective that propagates gradients through recurrent latent computation. SWITCH consistently outperforms prior hidden-state-recurrence latent reasoning approaches at similar scale. Mechanistic analysis through the boundary tokens further reveals three findings: (i) <swi> is a sharply localised, learned switching policy rather than a stylistic artefact; (ii) the latent step it opens performs problem-specific, causally important computation rather than acting as an inert placeholder; and (iii) that computation is concentrated at a single hidden-state transition on entry. Together, these results show that hidden-state-recurrence latent reasoning is both RL-trainable and open to direct mechanistic analysis, including of how on-policy RL itself improves the model from the inside.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis