Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
AuthorsQiangqiang He, Zhongheng Wu, ZiJian Wang
Resources
The paper argues that RL should give less credit to highly counterfactually sensitive tokens and shows that this improves long-chain-of-thought reasoning.
Key results
Fixed on-policy trajectories used to test counterfactual likelihood shifts.
Maximum reported proportion of tokens insignificant under both privileged conditions.
Mean Counterfactual Perturbation Consistency across the 400 trajectories.
Average CSCR improvement over the strongest competing method.
Average CSCR improvement over the strongest competing method.
Stable downweighting strength identified by ablation.
What the paper found
This paper challenges the assumption that privileged likelihood shifts provide reliable token-level supervision for long-chain-of-thought reinforcement learning with verifiable rewards. In 400 fixed on-policy trajectories from DAPO-17K, positive and negative counterfactual prompts produced sparse changes—up to 84.7% of tokens were insignificant under both conditions—and their full-vocabulary optimization signals substantially overlapped, with mean Counterfactual Perturbation Consistency of 0.583. Large shifts concentrated on replaceable discourse tokens such as “The,” “But,” and “Therefore,” while digits and mathematical operators carrying problem-specific reasoning were often insensitive. The proposed Counterfactual Sensitivity Credit Reallocation, or CSCR, therefore treats shift magnitude as sensitivity rather than learning value: it exponentially downweights highly sensitive tokens, preserves the verifier-defined GRPO direction, and renormalizes advantages to retain the trajectory’s total credit. On Qwen3-1.7B and Qwen3-4B, evaluated with Mean@32 across AMC23, AIME24, AIME25, AIME26, and SMT25, CSCR was best across all ten model–benchmark combinations, improving over the strongest competing method by 3.9 points and 1.7 points respectively. Ablations found that moderate downweighting, with attenuation strength 0.2, was stable, whereas trusting privileged shift signs caused rapid repetition and optimization collapse. The experiments use NVIDIA H800 GPUs, and robustness tests include reference solutions generated by DeepSeek-V4-Pro, positioning CSCR as a lightweight alternative to OPSD and other self-distillation methods.
Original abstract
Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.