Less is More: Early Stopping Rollout for On-Policy Distillation
AuthorsZhou Ziheng, Jiaqi Li, Huacong Tang, Ying Nian Wu, Demetri Terzopoulos
Resources
This work shows that for on-policy distillation, letting the student only roll out the first few response tokens can actually outperform full rollouts, while being faster and more stable to train.
Key results
ESR on Qwen2.5-Math-1.5B → Qwen3-1.7B
Baseline full-rollout on Qwen2.5-Math-1.5B → Qwen3-1.7B
ESR on Qwen2.5-Math-1.5B → Qwen3-1.7B
ESR versus full-rollout OPD per training step
ESR peak training memory in GB
Cascading Alignment effect on untrained late tokens
What the paper found
This paper from UCLA and the Beijing Institute of General Artificial Intelligence argues that on-policy distillation can fail because late rollout tokens become off-policy to the teacher, causing “Off-policy Teacher Decay.” The proposed fix, Early Stopping Rollout (ESR), is a one-line change: stop student rollouts after the first N response tokens and compute reverse-KL loss only on that early window. Across MATH-500, HumanEval, and BFCL, ESR consistently matches or beats full-rollout OPD across model families and scales, including Qwen2.5, Qwen3, and Gemma, with student sizes from 1.5B to 32B and teacher sizes from 1.7B to 72B. The method is especially strong in cross-generation and cross-family transfers, where full-rollout training often collapses. On MATH-500, ESR improves Qwen2.5-Math-1.5B → Qwen3-1.7B to 65.85 avg@4 versus 62.35 for OPD, and on HumanEval it reaches 42.10 pass@1 versus 40.20. It also cuts per-step wall-clock from about 194 s to 8 s, a 24× speedup, while reducing peak memory from 63.3 GB to 24.1 GB, about 2.6× lower. Mechanistically, the authors find a Cascading Alignment effect, where training only the first 100 tokens lowers KL by 30–40% even on untrained later positions, and a Sub-mode Commitment effect, where reverse-KL distillation steers students toward a supported teacher sub-mode and can even surpass the teacher.
Original abstract
On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rollouts with a teacher model. However, we observe ``Off-policy Teacher Decay'' problem in this paradigm: for the later tokens, with student's earlier trajectory as context that is off-policy to the teacher, the teacher's ability to produce a corrective score would decay, and may fall back to token-completion behavior learned in the pre-training stage. We empirically verify this problem, and we propose Early Stopping Rollout (ESR) to fix it: a simple yet effective distillation strategy that simply restricts the rollout generation to the first response tokens. We show that ESR both surpasses the full rollout OPD performance across model size, family, tasks and training regime, and exhibit much higher GPU efficiency and training stability, especially under cross model family scenarios. We further investigate the mechanism behind this surprising performance and discovered "Cascading Alignment" and "Sub-mode Commitment" effect of ESR that may explain why it works effectively and even sometimes exceeding the teacher model performance. Besides, we show that this position-based token selection strategy cannot be fully explainable by KL divergence and entropy signals.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.