GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards
AuthorsTej Deep Pala, Vernon Toh, Soujanya Poria
Resources
GRAIL improves LLM reasoning training by redistributing reinforcement learning credit to the tokens that matter most, boosting accuracy without needing step-by-step reward labels.
Key results
GRAIL vs GRPO across five models
GRAIL vs GRPO across five models
GRAIL result, up from 36.67% with GRPO
GRAIL result, up from 2.22% with GRPO
GRAIL performance compared with OAR-G's 55.75%
Wilcoxon signed-rank test on paired average accuracy scores
What the paper found
GRAIL, from DeCLaRe Lab at Nanyang Technological University, is a token-wise credit-assignment method for reinforcement learning with verifiable rewards that replaces GRPO’s uniform sequence-level advantage with gradient-activation saliency. Instead of paying for a process reward model, it backpropagates the final answer loss to the input embeddings, computes a first-order saliency score per token, log-scales and standardizes those scores, then clips them into a bounded reweighting scheme before the policy-gradient update. Across five models from the Qwen3, DeepSeek-R1-Distill-Llama-8B, and OctoThinker families trained on DeepMath-103K, GRAIL improves average accuracy by 3.60% and Pass@3 by 3.05% over GRPO on six math benchmarks, with especially large gains on AIME 2024: Qwen3-8B rises from 36.67% to 47.78% accuracy, while OctoThinker-8B rises from 2.22% to 4.44%. The method also beats OAR-G, reaching 62.04% average accuracy on Qwen3-8B versus 55.75%. A Wilcoxon signed-rank test reports W = 459.5 with p = 1.3 × 10−8, and the paper notes a 50% to 60% training-time overhead on 4 × NVIDIA H200 GPUs, trading compute for finer-grained reasoning alignment without step-level supervision.
Original abstract
Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually broadcast one sequence-level advantage to all tokens, or use costly process reward models (PRMs) for step-level supervision. Uniform advantage distribution assumes that all tokens contribute equally to the final reward. This dilutes the gradient signal, since flawed reasoning steps and filler words are updated as strongly as valid logical inferences. To address this, we introduce Gradient-Reweighted Advantage (GRAIL), an intrinsic token-wise advantage reweighting method. GRAIL uses gradient-activation saliency to place more weight on tokens that are more locally sensitive to the final answer. Evaluations across five models from the Qwen3, R1-distilled and OctoThinker families show that GRAIL consistently outperforms GRPO. GRAIL achieved an average improvement of 3.60% in accuracy and 3.05% in Pass@3, demonstrating that fine-grained reasoning alignment can be achieved without process-level supervision.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.