NTH

Discretizing Reward Models

AuthorsVijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao

June 26, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that reward models are often too fine-grained for their own good, and shows that discretizing their outputs can reduce reward hacking and produce better RL policies.

Key results

70.8
RewardBench 2 Ties score, Skywork V1 raw

Average of discriminative ability and specificity under the paper’s new metrics

74.9
RewardBench 2 Ties score, Skywork V1 clustered

Average of discriminative ability and specificity after reward clustering

69.2
RewardBench 2 Ties score, GRM raw

Average of discriminative ability and specificity under the paper’s new metrics

80.6
RewardBench 2 Ties score, GRM clustered

Average of discriminative ability and specificity after reward clustering

15%
Runtime overhead

Average GRPO training runtime increase from reward clustering

4
MC dropout samples

Default number of stochastic forward passes used to estimate reward variance

What the paper found

Discretizing Reward Models, from Carnegie Mellon University and Meta Superintelligence Labs, argues that continuous reward models can be dangerously oversensitive: they often assign different scores to responses that are equally useful, which creates exploitable reward gradients in RLHF. The paper separates reward quality into discriminative ability and specificity, showing that popular models such as Skywork V1, Skywork V2, GRM, and ArmoRM can score highly on standard benchmarks yet still fail this tie-sensitive criterion. The core result is that a training-free discretization method, reward clustering, uses Monte Carlo dropout to estimate reward uncertainty, hierarchically clusters response scores, and maps clusters to ordinal bins; in the theory, an optimal midpoint threshold can preserve perfect discriminative ability while eliminating oversensitivity in the binary case. On RewardBench 2’s Ties subset, clustered rewards improve the average of discriminative ability and specificity for all four reward models, with Skywork V1 rising from 70.8 to 74.9 and GRM from 69.2 to 80.6. In a mixed-reward IFEval simulation, discretization suppresses overoptimization of spurious hedging rewards, and in a verifier-free multi-task RL setup on 30K prompts each from RLVR-IFeval, RLVR-MATH, RLVR-GSM, and WildChat, discretized rewards were never significantly worse than raw rewards and often much better. The method adds only a modest 15% runtime overhead during GRPO training, while using 8 H100 GPUs and 4 dropout samples.

Original abstract

Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Reward models offer a tempting promise: they automatically estimate response quality in the absence of verifiers or human judges. Unlike "verifiable rewards" which typically produce binary scores, reward models typically produce continuous scores, allowing them to be sensitive to fine-grained differences in responses. However, we show this apparent strength is a serious weakness: many popular reward models are oversensitive, assigning different scores to equally good responses. Theoretically, we show that seemingly perfect reward models can be highly oversensitive; empirically, this oversensitivity can lead to bad policies. In place of existing notions of "reward model accuracy," we propose evaluating reward models using distinct measures of "discriminative ability" and "specificity" (the complement of oversensitivity). As a solution, we describe a training-free algorithm that uses Monte Carlo dropout on any neural reward model to produce discrete reward clusters. Theoretically, we prove there exist discretizations that reduce oversensitivity at minimal expense of discriminative ability; empirically we show, in both controlled and natural RL settings, that discretizing rewards leads to less reward hacking and better policies than training on the original rewards.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →