Rethinking On-Policy Distillation of Large Language Models II: One Training Example
AuthorsZixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao
Resources
This study finds that on-policy distillation of language models may need far fewer examples than expected, but many more training steps to fully absorb the supervision.
Key results
Fraction of the semantic state space reached by full-data OPD rollouts.
Coverage achieved by 16 semantically diverse queries.
Full-data OPD gain recovered after 1000 steps, with 68.4 versus 72.1 average accuracy.
Average validation accuracy reached by full-data multi-teacher OPD.
Average validation accuracy from 16 queries per domain.
What the paper found
This paper tests on-policy distillation at its data-minimal limit: training a student on one query while the student generates rollouts and a teacher supplies dense token-level distributions at every visited prefix. Across DeepSeek-R1-Distill-Qwen-1.5B, Llama-3B-It, OLMo-7B-It-DPO, and Qwen-Coder-1.5B, one-shot OPD continues improving for hundreds of steps across mathematics, code, instruction following, and tool use. On DAPO-Math-17K, it reaches 68.4 average accuracy after 1000 steps versus 72.1 for full-data OPD, recovering 72% of the full-data gain across MATH-500, AMC 2023, and AIME 2025. The proposed explanation is that OPD is “data-overfed but algorithm-starved”: repeated rollouts from one query already reach 71.5% of the semantic state clusters visited by full-data training, while the student’s absorption rate for teacher supervision keeps declining regardless of query count. Selecting semantically diverse queries expands coverage; 16 queries reach 98.9% and match full-data OPD. The same pattern holds in multi-teacher OPD: 16 queries per domain reach 52.9 average validation accuracy versus 52.8 for full-data MOPD across math, code, and instruction following. Even content-light templates and off-domain WildChat prompts, derived from ChatGPT interaction logs, can approach real-query performance, suggesting that an input’s main role is to induce useful reasoning states rather than state a high-quality task. The findings redirect optimization toward faster supervision absorption and state-aware data selection, with implications for post-training systems built around Qwen3, DeepSeek, and other frontier models.
Original abstract
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.