NTH

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

AuthorsXiaobo Wang, Tong Wu, Min Tang, Jiaqi Li, Qi Liu, Zilong Zheng

June 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces SAVE, a method that lets reward models learn from a policy’s own on-policy outputs using value-guided self-supervision, improving alignment without relying only on fresh human labels.

Key results

76.0
Reward model average

baseline average accuracy before SAVE

77.3
Reward model average

SAVE average accuracy with Qwen3-4B-Instruct-2507

51.68%
AlpacaEval 2 LC win rate

baseline policy performance before using the improved reward model

54.24%
AlpacaEval 2 LC win rate

policy performance with the improved reward model

33.9%
Arena-Hard-v2.0 win rate

policy performance with the improved reward model

What the paper found

SAVE, from researchers at the University of Science and Technology of China and BIGAI, is a self-supervised RLHF method that turns the policy’s own on-policy rollouts into training signal for the reward model instead of relying on additional human labels or an external judge. The core idea is to add a prompt-specific value head to the reward model, use the value estimate as an adaptive anchor, filter out near-zero-advantage responses with a curriculum threshold, and update the reward head with a contrastive objective on positive versus negative on-policy feedback. The paper formalizes this as a reward-model-centric minimax problem and shows that, under local conditions, standard policy optimization on KL-regularized objectives naturally generates informative samples where the current reward model is most miscalibrated. On six benchmarks—RewardBench, RewardBench 2, RM-Bench, PPE Preference, PPE Correctness, and JudgeBench—SAVE raises the average reward-model accuracy from 76.0 to 77.3 with the stronger Qwen3-4B-Instruct-2507 policy, and from 76.0 to 76.7 with Qwen2.5-3B-Instruct. The improved reward model then transfers to downstream RLHF, increasing AlpacaEval 2 length-controlled win rate from 51.68% to 54.24% and Arena-Hard-v2.0 win rate from 30.2% to 33.9%.

Original abstract

Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models. It is dramatically worse as the policy evolves beyond the static RM training. Therefore, we propose SAVE (Self-supervised reward model improvement via Value-Anchored On-policy feedback), a framework that grades on-policy responses as feedback by using the value function for on-policy RM training. SAVE naturally converts the reward-graded on-policy responses into supervision with a prompt-specific value head as an adaptive anchor. It computes RM advantages and filters ambiguous samples to update the RM via a contrastive objective. The effectiveness of SAVE for enhancing RM training is strongly validated through rigorous empirical evaluation across six diverse benchmarks. It achieves outperforming results across all datasets while maintaining consistent improvements across three RL algorithms (GRPO, RLOO, GSPO) and different policy backbones.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →