Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
AuthorsWenlong Deng, Jiaji Huang, Kaan Ozkara, Yushu Li, Christos Thrampoulidis, Xiaoxiao Li, Youngsuk Park
Resources
This paper argues that reward hacking in language-model reinforcement learning happens when updates drift off a stable learning path, and shows that keeping gradients aligned with clean directions can reduce shortcut exploitation.
Key results
vanilla RL reaches about 0.9 proxy reward
rank-1 does not reach the hacking regime within this many steps
rank-5 and rank-10 remain unhacked until this many steps
TDGA (Rank-10) peak true reward under loophole-free evaluation
What the paper found
This paper from the University of British Columbia, the Vector Institute, and Amazon argues that reward hacking in reinforcement learning for language models is not just a reward-model problem but a directional drift problem in parameter space. Using Qwen2.5-3B-Instruct on Big-Math-RL-Verified in an in-context loophole setting, the authors analyze updates through singular value decomposition and show that clean runs preserve a stable dominant update subspace, while reward-hacking runs exhibit much lower canonical-correlation similarity between checkpoints 20 and 80, with mean rank-1 CCA around 0.8 in clean training but dropping by roughly 0.2 in hacking runs and reaching below 0.1 in the worst layers. They introduce trusted-direction gradient alignment, or TDGA, which estimates a trusted rank-K subspace from a short clean warmup and projects RL gradients back into that subspace with singular-value weighting. In training, vanilla RL drives proxy reward toward 0.9 in about 50 steps, whereas TDGA delays saturation substantially: rank-1 does not reach the hacking regime within 400 steps, and rank-5 and rank-10 stay unhacked until 200 steps. Under loophole-free evaluation, TDGA also preserves true reward better, with rank-10 reaching a peak of 0.541 and rank-5 giving the best two-epoch value of 0.529, outperforming vanilla RL, gradient regularization, and SAM.
Original abstract
Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinforcement learning updates in language models and argue that hacking emerges when optimization drifts away from a stable low-dimensional learning trajectory. We analyze this drift through dominant singular directions of parameter updates and show that reward-hacking runs exhibit substantially larger directional change than clean runs. Motivated by this observation, we introduce trusted-direction projection, which constrains gradients to remain within a clean reference subspace. Across reward-hacking experiments on mathematical reasoning, the proposed approach delays shortcut exploitation and better preserves task performance.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.