NTH

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

AuthorsJing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng

July 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that LLM RL should optimize what actually improves inference-time behavior, not just training-time policy updates, and introduces a method to make that happen more reliably.

Key results

66.71%
Qwen3-4B Avg pass@1

Best average reasoning score for MIPU under FP8-quantized rollout on Qwen3-4B

53.97%
Qwen3-1.7B Avg pass@1

Best average reasoning score for MIPU under FP8-quantized rollout on Qwen3-1.7B

What the paper found

This paper argues that the real objective in LLM reinforcement learning is not optimizing the training policy inside the trainer, but achieving monotonic improvement of the inference policy used for rollout and deployment. The authors formalize this as Monotonic Inference Policy Improvement, or MIPI, and implement it with Monotonic Inference Policy Update, MIPU, a two-step framework: Step 1 constructs sampler-referenced candidate updates with truncated importance correction, and Step 2 validates the synchronized candidate using a post-update inference-gap proxy before accepting or rolling back the checkpoint. The method is tested under FP8-quantized rollout, a deliberately high-mismatch setting, on Alibaba’s Qwen3-4B and Qwen3-1.7B models trained with GRPO-style RL on DAPO-Math-17 and DeepMath-103K. On five math reasoning benchmarks—MATH-500, AIME24, AMC23, Minerva, and OlympiadBench—MIPU achieves the best average pass@1 on both scales, reaching 66.71% on Qwen3-4B and 53.97% on Qwen3-1.7B, while also avoiding the collapse and sharp degradation seen in GRPO, MIS, and LR-decay baselines. Ablations show the two steps are complementary: Step 1 improves candidate quality, Step 2 filters unreliable synchronized updates, and a random rollback control still collapses despite rejecting more updates, confirming that the inference-gap signal—not mere conservatism—is what stabilizes training.

Original abstract

Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: LLM adopts separate inference and training engines for generation efficiency and training precision, which in practice exhibits inconsistent probabilities for the same trajectories on training and inference sides, even with synchronized model parameters. This naturally induces a special type of off-policyness ever existing and poisoning the training. Prior works have made various efforts in addressing the off-policyness to stabilize the training policies under the mismatch. In this paper, we point out the objective misalignment neglected by existing works that an effective update to the policy in the training engine not necessarily ensures the improvement of the inference policy, i.e., the one used in deployment. To this end, we propose a new policy optimization objective for LLM RL, named Monotonic Inference Policy Improvement (MIPI). Following this principle, we introduce Monotonic Inference Policy Update (MIPU), a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy. Experiments conducted on two model scales under high mismatch show that MIPU improves average reasoning performance and training stability.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →