MInTRL: Off-policy Intervention can boost On-policy RL
AuthorsMingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong
AffiliationsAmazon Web Services · Boston University
Resources
MInTRL helps reinforcement-learning systems discover better solutions by making small, targeted corrections during otherwise on-policy rollouts.
Key results
MInTRL-Const average across AIME 2025, AIME 2026, and HMMT February 2025.
MInTRL-Const average across LiveCodeBench, HumanEval+, and MBPP+.
MInTRL-Const average across the three math benchmarks.
MInTRL-Const average across the three code benchmarks.
Performance peaks within the reported roughly 2–4% off-policy intervention-token range.
Qwen3-1.7B mathematics wall-clock speedup during the intervention phase.
What the paper found
MInTRL, or Minimal Intervention Reinforcement Learning, addresses a central weakness of reinforcement learning with verifiable rewards: purely on-policy sampling cannot reliably discover reasoning paths outside a model’s finite-sampling frontier, while full off-policy trajectories create distribution shift. During rollout, the current policy generates most tokens, a judge detects the earliest incorrect math or code step, and an intervention policy replaces only the erroneous suffix with a short correction before control returns to the student. Training uses sequence-level advantage regression, with on-policy control rollouts providing the value baseline, so mixed-provenance trajectories are learned without behavior-policy importance sampling or direct imitation of the teacher. Experiments with Qwen3-1.7B and Qwen3-4B on AceReason-Nemotron, DeepCoder-Preview, AIME 2025, AIME 2026, HMMT February 2025, LiveCodeBench, HumanEval+, and MBPP+ show that Qwen3-1.7B reaches mathematics and code averages of 35.45 and 61.95, while Qwen3-4B reaches 55.73 and 72.63 with MInTRL-Const. The method remains effective with self-intervention and with DeepSeek-V4-Flash as the judge, but its gains follow an inverted-U curve: performance peaks when intervention tokens comprise roughly 2–4%, whereas excessive intervention becomes harder to learn from. Against MENTOR, MInTRL-Const achieves up to 3.45× speedup, despite the extra judging overhead, because it avoids continuous token-level coordination. The results support a coverage–learnability trade-off in which sparse off-policy corrections expand exploration while preserving predominantly on-policy training.
Original abstract
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's output, replaces erroneous suffixes with short corrections, and immediately returns control to the policy. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. We show that sparse, local interventions can substantially improve coverage beyond finite-budget on-policy sampling while preserving the overall on-policy nature of the resulting trajectories. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines. Ablations show that MInTRL remains effective with self-intervention and across different judge policies, while performance peaks at moderate intervention intensity, highlighting the importance of intervening minimally. These results establish minimal intervention as an effective paradigm for enhancing on-policy RL.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.