The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
AuthorsXiaobo Wang, Tong Wu, Min Tang, Jiaqi Li, Qi Liu, Zilong Zheng
Resources
This paper introduces SAVE, a method that lets reward models learn from a policy’s own on-policy outputs using value-guided self-supervision, improving alignment without relying only on fresh human labels.
Key results
baseline average accuracy before SAVE
SAVE average accuracy with Qwen3-4B-Instruct-2507
baseline policy performance before using the improved reward model
policy performance with the improved reward model
policy performance with the improved reward model
What the paper found
SAVE, from researchers at the University of Science and Technology of China and BIGAI, is a self-supervised RLHF method that turns the policy’s own on-policy rollouts into training signal for the reward model instead of relying on additional human labels or an external judge. The core idea is to add a prompt-specific value head to the reward model, use the value estimate as an adaptive anchor, filter out near-zero-advantage responses with a curriculum threshold, and update the reward head with a contrastive objective on positive versus negative on-policy feedback. The paper formalizes this as a reward-model-centric minimax problem and shows that, under local conditions, standard policy optimization on KL-regularized objectives naturally generates informative samples where the current reward model is most miscalibrated. On six benchmarks—RewardBench, RewardBench 2, RM-Bench, PPE Preference, PPE Correctness, and JudgeBench—SAVE raises the average reward-model accuracy from 76.0 to 77.3 with the stronger Qwen3-4B-Instruct-2507 policy, and from 76.0 to 76.7 with Qwen2.5-3B-Instruct. The improved reward model then transfers to downstream RLHF, increasing AlpacaEval 2 length-controlled win rate from 51.68% to 54.24% and Arena-Hard-v2.0 win rate from 30.2% to 33.9%.
Original abstract
Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models. It is dramatically worse as the policy evolves beyond the static RM training. Therefore, we propose SAVE (Self-supervised reward model improvement via Value-Anchored On-policy feedback), a framework that grades on-policy responses as feedback by using the value function for on-policy RM training. SAVE naturally converts the reward-graded on-policy responses into supervision with a prompt-specific value head as an adaptive anchor. It computes RM advantages and filters ambiguous samples to update the RM via a contrastive objective. The effectiveness of SAVE for enhancing RM training is strongly validated through rigorous empirical evaluation across six diverse benchmarks. It achieves outperforming results across all datasets while maintaining consistent improvements across three RL algorithms (GRPO, RLOO, GSPO) and different policy backbones.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.