NTH

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

AuthorsQiangqiang He, Zhongheng Wu, ZiJian Wang

August 9, 2026 2 min read
Watch on YouTube
The one-line take

The paper argues that RL should give less credit to highly counterfactually sensitive tokens and shows that this improves long-chain-of-thought reasoning.

Key results

400
Diagnostic trajectories

Fixed on-policy trajectories used to test counterfactual likelihood shifts.

84.7%
Jointly insignificant-token upper bound

Maximum reported proportion of tokens insignificant under both privileged conditions.

0.583
Mean CPC

Mean Counterfactual Perturbation Consistency across the 400 trajectories.

3.9
Qwen3-1.7B gain

Average CSCR improvement over the strongest competing method.

1.7
Qwen3-4B gain

Average CSCR improvement over the strongest competing method.

0.2
Moderate attenuation strength

Stable downweighting strength identified by ablation.

What the paper found

This paper challenges the assumption that privileged likelihood shifts provide reliable token-level supervision for long-chain-of-thought reinforcement learning with verifiable rewards. In 400 fixed on-policy trajectories from DAPO-17K, positive and negative counterfactual prompts produced sparse changes—up to 84.7% of tokens were insignificant under both conditions—and their full-vocabulary optimization signals substantially overlapped, with mean Counterfactual Perturbation Consistency of 0.583. Large shifts concentrated on replaceable discourse tokens such as “The,” “But,” and “Therefore,” while digits and mathematical operators carrying problem-specific reasoning were often insensitive. The proposed Counterfactual Sensitivity Credit Reallocation, or CSCR, therefore treats shift magnitude as sensitivity rather than learning value: it exponentially downweights highly sensitive tokens, preserves the verifier-defined GRPO direction, and renormalizes advantages to retain the trajectory’s total credit. On Qwen3-1.7B and Qwen3-4B, evaluated with Mean@32 across AMC23, AIME24, AIME25, AIME26, and SMT25, CSCR was best across all ten model–benchmark combinations, improving over the strongest competing method by 3.9 points and 1.7 points respectively. Ablations found that moderate downweighting, with attenuation strength 0.2, was stable, whereas trusting privileged shift signs caused rapid repetition and optimization collapse. The experiments use NVIDIA H800 GPUs, and robustness tests include reference solutions generated by DeepSeek-V4-Pro, positioning CSCR as a lightweight alternative to OPSD and other self-distillation methods.

Original abstract

Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →