NTH

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards

AuthorsTej Deep Pala, Vernon Toh, Soujanya Poria

June 26, 2026 2 min read
Watch on YouTube
The one-line take

GRAIL improves LLM reasoning training by redistributing reinforcement learning credit to the tokens that matter most, boosting accuracy without needing step-by-step reward labels.

Key results

3.60%
average accuracy gain

GRAIL vs GRPO across five models

3.05%
Pass@3 gain

GRAIL vs GRPO across five models

47.78%
AIME 2024 Qwen3-8B accuracy

GRAIL result, up from 36.67% with GRPO

4.44%
AIME 2024 OctoThinker-8B accuracy

GRAIL result, up from 2.22% with GRPO

62.04%
Qwen3-8B average accuracy vs OAR-G

GRAIL performance compared with OAR-G's 55.75%

1.3e-8
statistical test p-value

Wilcoxon signed-rank test on paired average accuracy scores

What the paper found

GRAIL, from DeCLaRe Lab at Nanyang Technological University, is a token-wise credit-assignment method for reinforcement learning with verifiable rewards that replaces GRPO’s uniform sequence-level advantage with gradient-activation saliency. Instead of paying for a process reward model, it backpropagates the final answer loss to the input embeddings, computes a first-order saliency score per token, log-scales and standardizes those scores, then clips them into a bounded reweighting scheme before the policy-gradient update. Across five models from the Qwen3, DeepSeek-R1-Distill-Llama-8B, and OctoThinker families trained on DeepMath-103K, GRAIL improves average accuracy by 3.60% and Pass@3 by 3.05% over GRPO on six math benchmarks, with especially large gains on AIME 2024: Qwen3-8B rises from 36.67% to 47.78% accuracy, while OctoThinker-8B rises from 2.22% to 4.44%. The method also beats OAR-G, reaching 62.04% average accuracy on Qwen3-8B versus 55.75%. A Wilcoxon signed-rank test reports W = 459.5 with p = 1.3 × 10−8, and the paper notes a 50% to 60% training-time overhead on 4 × NVIDIA H200 GPUs, trading compute for finer-grained reasoning alignment without step-level supervision.

Original abstract

Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually broadcast one sequence-level advantage to all tokens, or use costly process reward models (PRMs) for step-level supervision. Uniform advantage distribution assumes that all tokens contribute equally to the final reward. This dilutes the gradient signal, since flawed reasoning steps and filler words are updated as strongly as valid logical inferences. To address this, we introduce Gradient-Reweighted Advantage (GRAIL), an intrinsic token-wise advantage reweighting method. GRAIL uses gradient-activation saliency to place more weight on tokens that are more locally sensitive to the final answer. Evaluations across five models from the Qwen3, R1-distilled and OctoThinker families show that GRAIL consistently outperforms GRPO. GRAIL achieved an average improvement of 3.60% in accuracy and 3.05% in Pass@3, demonstrating that fine-grained reasoning alignment can be achieved without process-level supervision.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →