NTH

Less is More: Early Stopping Rollout for On-Policy Distillation

AuthorsZhou Ziheng, Jiaqi Li, Huacong Tang, Ying Nian Wu, Demetri Terzopoulos

June 13, 2026 2 min read
Watch on YouTube
The one-line take

This work shows that for on-policy distillation, letting the student only roll out the first few response tokens can actually outperform full rollouts, while being faster and more stable to train.

Key results

65.85
MATH-500 avg@4

ESR on Qwen2.5-Math-1.5B → Qwen3-1.7B

62.35
OPD avg@4

Baseline full-rollout on Qwen2.5-Math-1.5B → Qwen3-1.7B

42.10
HumanEval pass@1

ESR on Qwen2.5-Math-1.5B → Qwen3-1.7B

24x
Wall-clock speedup

ESR versus full-rollout OPD per training step

24.1
Peak GPU memory

ESR peak training memory in GB

30-40%
KL drop

Cascading Alignment effect on untrained late tokens

What the paper found

This paper from UCLA and the Beijing Institute of General Artificial Intelligence argues that on-policy distillation can fail because late rollout tokens become off-policy to the teacher, causing “Off-policy Teacher Decay.” The proposed fix, Early Stopping Rollout (ESR), is a one-line change: stop student rollouts after the first N response tokens and compute reverse-KL loss only on that early window. Across MATH-500, HumanEval, and BFCL, ESR consistently matches or beats full-rollout OPD across model families and scales, including Qwen2.5, Qwen3, and Gemma, with student sizes from 1.5B to 32B and teacher sizes from 1.7B to 72B. The method is especially strong in cross-generation and cross-family transfers, where full-rollout training often collapses. On MATH-500, ESR improves Qwen2.5-Math-1.5B → Qwen3-1.7B to 65.85 avg@4 versus 62.35 for OPD, and on HumanEval it reaches 42.10 pass@1 versus 40.20. It also cuts per-step wall-clock from about 194 s to 8 s, a 24× speedup, while reducing peak memory from 63.3 GB to 24.1 GB, about 2.6× lower. Mechanistically, the authors find a Cascading Alignment effect, where training only the first 100 tokens lowers KL by 30–40% even on untrained later positions, and a Sub-mode Commitment effect, where reverse-KL distillation steers students toward a supported teacher sub-mode and can even surpass the teacher.

Original abstract

On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rollouts with a teacher model. However, we observe ``Off-policy Teacher Decay'' problem in this paradigm: for the later tokens, with student's earlier trajectory as context that is off-policy to the teacher, the teacher's ability to produce a corrective score would decay, and may fall back to token-completion behavior learned in the pre-training stage. We empirically verify this problem, and we propose Early Stopping Rollout (ESR) to fix it: a simple yet effective distillation strategy that simply restricts the rollout generation to the first response tokens. We show that ESR both surpasses the full rollout OPD performance across model size, family, tasks and training regime, and exhibit much higher GPU efficiency and training stability, especially under cross model family scenarios. We further investigate the mechanism behind this surprising performance and discovered "Cascading Alignment" and "Sub-mode Commitment" effect of ESR that may explain why it works effectively and even sometimes exceeding the teacher model performance. Besides, we show that this position-based token selection strategy cannot be fully explainable by KL divergence and entropy signals.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis