STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
AuthorsHaipeng Luo, Qingfeng Sun, Songli Wu, Can Xu, Wenfeng Deng, Han Hu, Yansong Tang
STARE is a new training trick for LLM reasoning that keeps exploration alive by reweighting the tokens most responsible for entropy collapse during reinforcement learning.
Key results
STARE is evaluated across LLM scales from 1.5B to 32B
Closed-loop gating activates when batch entropy falls below Htgt = 0.3
STARE outperforms DAPO and other baselines on AIME24 and AIME25 by 4%–8% in average accuracy
Main experiments use a batch-internal top-P% surprisal proxy with P = 10%
What the paper found
STARE, from Tsinghua University and Tencent Hunyuan, targets a specific failure mode of GRPO-style reinforcement learning with verifiable rewards: policy entropy collapse that prematurely ends long-horizon training in large language models such as Qwen2.5, Qwen3, and DeepSeek-R1 variants. The paper’s core technical result is a first-order entropy analysis showing a token-level advantage-surprisal four-quadrant structure, where entropy change decomposes into the product of trajectory advantage and a local surprisal-sensitive term; this reveals that shared trajectory-level advantages systematically over-reinforce low-surprisal majority tokens while underweighting the high-surprisal minority that preserves exploration. STARE operationalizes this insight by selecting entropy-critical tokens via batch-internal top-P% surprisal quantiles, reweighting their effective advantages inside the clipped GRPO surrogate, and using a closed-loop target-entropy gate at Htgt = 0.3 to turn intervention on only when batch entropy falls below target. Across models from 1.5B to 32B, the method sustains stable RL for thousands of steps, including over 5k steps on 7B and over 1.5k steps on 14B and 32B, while maintaining entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other baselines by 4%–8% in average accuracy, with the strongest reported single-setting gains reaching 54.4%, 66.3%, and 60.4% average accuracy in the Short CoT, Long CoT, and tool-use regimes, respectively; ablations show that fixed W = 1.1, P = 10%, and batch-level gating are the most robust defaults.
Original abstract
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level advantage and an entropy sensitivity function over the next-token distribution, yielding an advantage-surprisal four-quadrant structure and a near-criticality property. Motivated by it, we propose STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability), which identifies entropy-critical token subsets via batch-internal surprisal quantiles, selectively reweights their effective advantages, and incorporates a target-entropy closed-loop gate for stable entropy regulation. Across model scales from 1.5B to 32B and three task families (Short CoT, Long CoT, and Multi-Turn Tool Use), STARE sustains stable RL training over thousands of steps while maintaining policy entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other competitive baselines by 4%-8% in average accuracy, with reflection tokens and response length growing in tandem, indicating sustained exploration-exploitation balance that further unlocks RL training potential.Code is available at https://github.com/hp-luo/STARE.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.