NTH

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

AuthorsZixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao

September 9, 2026 3 min read
Watch on YouTube
The one-line take

This study finds that on-policy distillation of language models may need far fewer examples than expected, but many more training steps to fully absorb the supervision.

Key results

71.5%
One-shot state coverage

Fraction of the semantic state space reached by full-data OPD rollouts.

98.9%
16-query state coverage

Coverage achieved by 16 semantically diverse queries.

72%
One-shot math gap recovery

Full-data OPD gain recovered after 1000 steps, with 68.4 versus 72.1 average accuracy.

52.8
Full-data MOPD accuracy

Average validation accuracy reached by full-data multi-teacher OPD.

52.9
16-shot MOPD accuracy

Average validation accuracy from 16 queries per domain.

What the paper found

This paper tests on-policy distillation at its data-minimal limit: training a student on one query while the student generates rollouts and a teacher supplies dense token-level distributions at every visited prefix. Across DeepSeek-R1-Distill-Qwen-1.5B, Llama-3B-It, OLMo-7B-It-DPO, and Qwen-Coder-1.5B, one-shot OPD continues improving for hundreds of steps across mathematics, code, instruction following, and tool use. On DAPO-Math-17K, it reaches 68.4 average accuracy after 1000 steps versus 72.1 for full-data OPD, recovering 72% of the full-data gain across MATH-500, AMC 2023, and AIME 2025. The proposed explanation is that OPD is “data-overfed but algorithm-starved”: repeated rollouts from one query already reach 71.5% of the semantic state clusters visited by full-data training, while the student’s absorption rate for teacher supervision keeps declining regardless of query count. Selecting semantically diverse queries expands coverage; 16 queries reach 98.9% and match full-data OPD. The same pattern holds in multi-teacher OPD: 16 queries per domain reach 52.9 average validation accuracy versus 52.8 for full-data MOPD across math, code, and instruction following. Even content-light templates and off-domain WildChat prompts, derived from ChatGPT interaction logs, can approach real-query performance, suggesting that an input’s main role is to induce useful reasoning states rather than state a high-quality task. The findings redirect optimization toward faster supervision absorption and state-aware data selection, with implications for post-training systems built around Qwen3, DeepSeek, and other frontier models.

Original abstract

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis