NTH

MInTRL: Off-policy Intervention can boost On-policy RL

AuthorsMingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong

AffiliationsAmazon Web Services · Boston University

September 20, 2026 2 min read
Watch on YouTube
The one-line take

MInTRL helps reinforcement-learning systems discover better solutions by making small, targeted corrections during otherwise on-policy rollouts.

Key results

35.45
Qwen3-1.7B mathematics average

MInTRL-Const average across AIME 2025, AIME 2026, and HMMT February 2025.

61.95
Qwen3-1.7B code average

MInTRL-Const average across LiveCodeBench, HumanEval+, and MBPP+.

55.73
Qwen3-4B mathematics average

MInTRL-Const average across the three math benchmarks.

72.63
Qwen3-4B code average

MInTRL-Const average across the three code benchmarks.

4%
Peak intervention intensity

Performance peaks within the reported roughly 2–4% off-policy intervention-token range.

3.45×
Maximum speedup over MENTOR

Qwen3-1.7B mathematics wall-clock speedup during the intervention phase.

What the paper found

MInTRL, or Minimal Intervention Reinforcement Learning, addresses a central weakness of reinforcement learning with verifiable rewards: purely on-policy sampling cannot reliably discover reasoning paths outside a model’s finite-sampling frontier, while full off-policy trajectories create distribution shift. During rollout, the current policy generates most tokens, a judge detects the earliest incorrect math or code step, and an intervention policy replaces only the erroneous suffix with a short correction before control returns to the student. Training uses sequence-level advantage regression, with on-policy control rollouts providing the value baseline, so mixed-provenance trajectories are learned without behavior-policy importance sampling or direct imitation of the teacher. Experiments with Qwen3-1.7B and Qwen3-4B on AceReason-Nemotron, DeepCoder-Preview, AIME 2025, AIME 2026, HMMT February 2025, LiveCodeBench, HumanEval+, and MBPP+ show that Qwen3-1.7B reaches mathematics and code averages of 35.45 and 61.95, while Qwen3-4B reaches 55.73 and 72.63 with MInTRL-Const. The method remains effective with self-intervention and with DeepSeek-V4-Flash as the judge, but its gains follow an inverted-U curve: performance peaks when intervention tokens comprise roughly 2–4%, whereas excessive intervention becomes harder to learn from. Against MENTOR, MInTRL-Const achieves up to 3.45× speedup, despite the extra judging overhead, because it avoids continuous token-level coordination. The results support a coverage–learnability trade-off in which sparse off-policy corrections expand exploration while preserving predominantly on-policy training.

Original abstract

Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's output, replaces erroneous suffixes with short corrections, and immediately returns control to the policy. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. We show that sparse, local interventions can substantially improve coverage beyond finite-budget on-policy sampling while preserving the overall on-policy nature of the resulting trajectories. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines. Ablations show that MInTRL remains effective with self-intervention and across different judge policies, while performance peaks at moderate intervention intensity, highlighting the importance of intervening minimally. These results establish minimal intervention as an effective paradigm for enhancing on-policy RL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →