Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
AuthorsRui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong
Resources
This paper explains why on-policy distillation can go wrong in LLM training and shows that regulating its guidance signals can make exploration more reliable.
Key results
OPD provides its clearest pass@k gains when k ≤ 64.
The student’s AIME 2025 accuracy remains near 2% with the strongest teacher.
Qwen3-1.7B-Base distilled from Qwen3-1.7B-GRPO achieves an overall average of 28.1.
Soft log-scale compression raises the same student’s overall average to 30.4.
What the paper found
Rui Wang and colleagues at The Chinese University of Hong Kong and Tencent AI Lab systematically analyze on-policy distillation, or OPD, using Qwen3-1.7B-Base students and Qwen3 teachers trained on the Nemotron-Cascade Math dataset across seven reasoning benchmarks. They find that OPD is primarily an exploration catalyst: dense token-level log-ratio guidance helps the student discover reasoning paths already within its capability, improving low-sample pass@k performance for k ≤ 64 without raising the asymptotic ceiling. Under fixed compute, one rollout per prompt, n=1, and broader prompt coverage outperform repeated sampling of the same problems. The paper identifies two pathologies: Student-Teacher Mismatch, where a large distributional gap makes a stronger teacher’s signal misleading, and Length Exploitation, where sequence-averaged advantages reward either filler-token padding or premature truncation. Notably, distillation from Qwen3-4B-GRPO leaves AIME 2025 accuracy near 2%, showing that teacher capability alone does not determine teaching quality. The authors regulate token advantages with hard clipping, ãt=clip(∆ℓt,cmin,cmax), or soft log-scale compression, ãt=sign(∆ℓt)·log(1+|∆ℓt|), requiring no additional off-policy compute. With Qwen3-1.7B-GRPO as teacher, log-scale compression raises the student’s overall average from 28.1 for naive OPD to 30.4, while regulated OPD can outperform methods using a massive 30B teacher, establishing signal fidelity and capacity matching—not brute-force teacher scale—as the central determinants of successful OPD.
Original abstract
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.