The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
AuthorsJing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
Resources
This paper argues that LLM RL should optimize what actually improves inference-time behavior, not just training-time policy updates, and introduces a method to make that happen more reliably.
Key results
Best average reasoning score for MIPU under FP8-quantized rollout on Qwen3-4B
Best average reasoning score for MIPU under FP8-quantized rollout on Qwen3-1.7B
What the paper found
This paper argues that the real objective in LLM reinforcement learning is not optimizing the training policy inside the trainer, but achieving monotonic improvement of the inference policy used for rollout and deployment. The authors formalize this as Monotonic Inference Policy Improvement, or MIPI, and implement it with Monotonic Inference Policy Update, MIPU, a two-step framework: Step 1 constructs sampler-referenced candidate updates with truncated importance correction, and Step 2 validates the synchronized candidate using a post-update inference-gap proxy before accepting or rolling back the checkpoint. The method is tested under FP8-quantized rollout, a deliberately high-mismatch setting, on Alibaba’s Qwen3-4B and Qwen3-1.7B models trained with GRPO-style RL on DAPO-Math-17 and DeepMath-103K. On five math reasoning benchmarks—MATH-500, AIME24, AMC23, Minerva, and OlympiadBench—MIPU achieves the best average pass@1 on both scales, reaching 66.71% on Qwen3-4B and 53.97% on Qwen3-1.7B, while also avoiding the collapse and sharp degradation seen in GRPO, MIS, and LR-decay baselines. Ablations show the two steps are complementary: Step 1 improves candidate quality, Step 2 filters unreliable synchronized updates, and a random rollback control still collapses despite rejecting more updates, confirming that the inference-gap signal—not mere conservatism—is what stabilizes training.
Original abstract
Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: LLM adopts separate inference and training engines for generation efficiency and training precision, which in practice exhibits inconsistent probabilities for the same trajectories on training and inference sides, even with synchronized model parameters. This naturally induces a special type of off-policyness ever existing and poisoning the training. Prior works have made various efforts in addressing the off-policyness to stabilize the training policies under the mismatch. In this paper, we point out the objective misalignment neglected by existing works that an effective update to the policy in the training engine not necessarily ensures the improvement of the inference policy, i.e., the one used in deployment. To this end, we propose a new policy optimization objective for LLM RL, named Monotonic Inference Policy Improvement (MIPI). Following this principle, we introduce Monotonic Inference Policy Update (MIPU), a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy. Experiments conducted on two model scales under high mismatch show that MIPU improves average reasoning performance and training stability.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.