Score Centering Stabilizes Off-policy Reinforcement Learning
AuthorsMartin Marek, Max Ryabinin
AffiliationsTogether AI
Resources
A simple score-centering correction makes off-policy RL for language models more stable when training and inference engines do not perfectly agree.
Key results
Qwen3-30B-A3B-Base on INTELLECT-2 math with FP8 weights and activations plus FP4 KV cache
Truncated importance sampling under the same Qwen3-30B-A3B-Base quantization setting
Qwen3-30B-A3B-Base on INTELLECT-2 math with INT8 weights and activations plus INT4 KV cache
Truncated importance sampling under the same severe quantization setting
Top-k approximation used to implement score centering
Approximate runtime increase of top-128 score centering versus baseline methods
What the paper found
This paper identifies training-inference mismatch as a central cause of instability in off-policy reinforcement learning for large language models, a problem relevant to systems behind models from OpenAI, DeepSeek, and Microsoft. The key mechanism is drift: when a quantized or stale sampler generates rollouts and a trainer computes gradients, the policy-gradient update contains a bias that repeatedly distills the trainer toward the sampler; synchronization then compounds the error until rewards collapse. The proposed solution, score centering, subtracts the sampler’s expected token score from each trainer score, canceling drift additively without importance ratios, clipping, masking, or a critic. Experiments spanning Qwen3 models from 0.6B to 30B parameters on Countdown and INTELLECT-2 show that score centering matches or beats truncated and masked importance sampling under severe quantization. With Qwen3-30B-A3B-Base on INTELLECT-2 math under FP8 weights and activations plus FP4 KV cache, score centering reaches 52% training accuracy versus 51% for truncated importance sampling; under INT8 weights and activations plus INT4 KV cache, it reaches 30% versus 12%. Under severe staleness, with the sampler updated every 64 steps, combining score centering with importance sampling performs best. For practical memory costs, the method logs only the sampler’s top 128 logprobs, matching full-vocabulary centering, with runs finishing within 1% of baseline wall-clock time.
Original abstract
Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive "score centering" correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.