NTH

Score Centering Stabilizes Off-policy Reinforcement Learning

AuthorsMartin Marek, Max Ryabinin

AffiliationsTogether AI

September 20, 2026 2 min read
Watch on YouTube
The one-line take

A simple score-centering correction makes off-policy RL for language models more stable when training and inference engines do not perfectly agree.

Key results

52%
FP4 KV score-centering accuracy

Qwen3-30B-A3B-Base on INTELLECT-2 math with FP8 weights and activations plus FP4 KV cache

51%
FP4 KV TIS accuracy

Truncated importance sampling under the same Qwen3-30B-A3B-Base quantization setting

30%
INT4 KV score-centering accuracy

Qwen3-30B-A3B-Base on INTELLECT-2 math with INT8 weights and activations plus INT4 KV cache

12%
INT4 KV TIS accuracy

Truncated importance sampling under the same severe quantization setting

128
Sampler logprob top-k

Top-k approximation used to implement score centering

1%
Wall-clock overhead

Approximate runtime increase of top-128 score centering versus baseline methods

What the paper found

This paper identifies training-inference mismatch as a central cause of instability in off-policy reinforcement learning for large language models, a problem relevant to systems behind models from OpenAI, DeepSeek, and Microsoft. The key mechanism is drift: when a quantized or stale sampler generates rollouts and a trainer computes gradients, the policy-gradient update contains a bias that repeatedly distills the trainer toward the sampler; synchronization then compounds the error until rewards collapse. The proposed solution, score centering, subtracts the sampler’s expected token score from each trainer score, canceling drift additively without importance ratios, clipping, masking, or a critic. Experiments spanning Qwen3 models from 0.6B to 30B parameters on Countdown and INTELLECT-2 show that score centering matches or beats truncated and masked importance sampling under severe quantization. With Qwen3-30B-A3B-Base on INTELLECT-2 math under FP8 weights and activations plus FP4 KV cache, score centering reaches 52% training accuracy versus 51% for truncated importance sampling; under INT8 weights and activations plus INT4 KV cache, it reaches 30% versus 12%. Under severe staleness, with the sampler updated every 64 steps, combining score centering with importance sampling performs best. For practical memory costs, the method logs only the sampler’s top 128 logprobs, matching full-vocabulary centering, with runs finishing within 1% of baseline wall-clock time.

Original abstract

Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive "score centering" correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →