NTH

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

AuthorsHaipeng Luo, Qingfeng Sun, Songli Wu, Can Xu, Wenfeng Deng, Han Hu, Yansong Tang

June 20, 2026 2 min read
Watch on YouTube
The one-line take

STARE is a new training trick for LLM reasoning that keeps exploration alive by reweighting the tokens most responsible for entropy collapse during reinforcement learning.

Key results

1.5B-32B
Model scale range

STARE is evaluated across LLM scales from 1.5B to 32B

0.3
Target entropy

Closed-loop gating activates when batch entropy falls below Htgt = 0.3

4-8%
Average accuracy gain over baselines

STARE outperforms DAPO and other baselines on AIME24 and AIME25 by 4%–8% in average accuracy

10%
Default top-surprisal ratio

Main experiments use a batch-internal top-P% surprisal proxy with P = 10%

What the paper found

STARE, from Tsinghua University and Tencent Hunyuan, targets a specific failure mode of GRPO-style reinforcement learning with verifiable rewards: policy entropy collapse that prematurely ends long-horizon training in large language models such as Qwen2.5, Qwen3, and DeepSeek-R1 variants. The paper’s core technical result is a first-order entropy analysis showing a token-level advantage-surprisal four-quadrant structure, where entropy change decomposes into the product of trajectory advantage and a local surprisal-sensitive term; this reveals that shared trajectory-level advantages systematically over-reinforce low-surprisal majority tokens while underweighting the high-surprisal minority that preserves exploration. STARE operationalizes this insight by selecting entropy-critical tokens via batch-internal top-P% surprisal quantiles, reweighting their effective advantages inside the clipped GRPO surrogate, and using a closed-loop target-entropy gate at Htgt = 0.3 to turn intervention on only when batch entropy falls below target. Across models from 1.5B to 32B, the method sustains stable RL for thousands of steps, including over 5k steps on 7B and over 1.5k steps on 14B and 32B, while maintaining entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other baselines by 4%–8% in average accuracy, with the strongest reported single-setting gains reaching 54.4%, 66.3%, and 60.4% average accuracy in the Short CoT, Long CoT, and tool-use regimes, respectively; ablations show that fixed W = 1.1, P = 10%, and batch-level gating are the most robust defaults.

Original abstract

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level advantage and an entropy sensitivity function over the next-token distribution, yielding an advantage-surprisal four-quadrant structure and a near-criticality property. Motivated by it, we propose STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability), which identifies entropy-critical token subsets via batch-internal surprisal quantiles, selectively reweights their effective advantages, and incorporates a target-entropy closed-loop gate for stable entropy regulation. Across model scales from 1.5B to 32B and three task families (Short CoT, Long CoT, and Multi-Turn Tool Use), STARE sustains stable RL training over thousands of steps while maintaining policy entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other competitive baselines by 4%-8% in average accuracy, with reflection tokens and response length growing in tandem, indicating sustained exploration-exploitation balance that further unlocks RL training potential.Code is available at https://github.com/hp-luo/STARE.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis