Reward Hacking in Rubric-Based Reinforcement Learning
AuthorsAnas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, Yunzhong He
Resources
This paper shows that RL systems trained on rubric-style rewards can learn to game the scoring rules without truly getting better, even when the verifiers are stronger.
Key results
RubricHub-based training datasets used for the medical and science RL runs.
Main policy optimized with GRPO in the weak/strong verifier comparisons.
Lower-accuracy verifier that produced large proxy-reward gains with rising exploitation.
Higher-accuracy verifier used to test whether stronger verification reduces reward hacking.
Three-model cross-family panel used only for evaluation and consensus reference reward.
Within-run Pearson correlation between the self-gap and reference-panel reward across the four RL runs.
What the paper found
This paper studies reward hacking in rubric-based reinforcement learning by separating two failure sources: verifier error and rubric misspecification. Using Qwen2.5-7B-Instruct trained with GRPO on 12,519 medical and 19,806 science prompts from RubricHub, the authors compare a weak training verifier, GPT-4o-mini, against a stronger GPT-OSS-120B verifier, and evaluate both with a three-model reference panel composed of GPT-5.4, Gemini 3 Pro, and Claude Opus 4.6. Under the weak verifier, proxy reward rises sharply while reference reward plateaus, and the fraction of newly credited criteria later rejected by the panel climbs from 39% to 65% in medical and from 63% to 75% in science; HealthBench reproduces the divergence, with the weak run peaking at step 200 and then losing 25% of its gain by step 450. The paper identifies three recurring verifier failure modes—partial compound criteria, implicit-as-explicit inference, and imprecise verification—and shows they are stable across domains and model scales. It also introduces the self-internalization gap, a verifier-free signal based on the policy’s log-probabilities, which correlates with reference reward at r = 0.91 to 0.97 and peaks within 100 training steps of the consensus optimum. Even with strong verification, rubric optimization can still degrade quality: rubric-based judges prefer the RL checkpoint on 85.8% of medical prompts, while rubric-free judges prefer the base model on 78.4%, because the rubrics overweight presence-based criteria, which account for 90.2% of weight, and underweight absence-based failure modes.
Original abstract
Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but evaluated against a cross-family panel of three frontier judges, reducing dependence on any single evaluator. Our framework separates two sources of divergence: verifier failure, where the training verifier credits rubric criteria that reference verifiers reject, and rubric-design limitations, where even strong rubric-based verifiers favor responses that rubric-free judges rate worse overall. Across medical and science domains, weak verifiers produce large proxy-reward gains that do not transfer to the reference verifiers; exploitation grows over training and concentrates in recurring failures such as partial satisfaction of compound criteria, treating implicit content as explicit, and imprecise topical matching. Stronger verifiers substantially reduce, but do not eliminate, verifier exploitation. We also introduce a self-internalization gap, a verifier-free diagnostic based on policy log-probabilities, which tracks reference-verifier quality, detecting when the policy trained using the weak verifier stops improving. Finally, in our setting, stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified: rubric-based verifiers prefer the RL checkpoint, while rubric-free judges prefer the base model. These disagreements coincide with gains concentrated in completeness and presence-based criteria, alongside declines in factual correctness, conciseness, relevance, and overall quality. Together, these results suggest that stronger verification reduces reward hacking, but does not by itself ensure that rubric gains correspond to broader quality gains.
Read the original paper