NTH

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

AuthorsRui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong

July 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains why on-policy distillation can go wrong in LLM training and shows that regulating its guidance signals can make exploration more reliable.

Key results

64
Low-sample exploration range

OPD provides its clearest pass@k gains when k ≤ 64.

2%
AIME 2025 accuracy under Qwen3-4B-GRPO guidance

The student’s AIME 2025 accuracy remains near 2% with the strongest teacher.

28.1
Naive OPD overall average

Qwen3-1.7B-Base distilled from Qwen3-1.7B-GRPO achieves an overall average of 28.1.

30.4
Log-scale OPD overall average

Soft log-scale compression raises the same student’s overall average to 30.4.

What the paper found

Rui Wang and colleagues at The Chinese University of Hong Kong and Tencent AI Lab systematically analyze on-policy distillation, or OPD, using Qwen3-1.7B-Base students and Qwen3 teachers trained on the Nemotron-Cascade Math dataset across seven reasoning benchmarks. They find that OPD is primarily an exploration catalyst: dense token-level log-ratio guidance helps the student discover reasoning paths already within its capability, improving low-sample pass@k performance for k ≤ 64 without raising the asymptotic ceiling. Under fixed compute, one rollout per prompt, n=1, and broader prompt coverage outperform repeated sampling of the same problems. The paper identifies two pathologies: Student-Teacher Mismatch, where a large distributional gap makes a stronger teacher’s signal misleading, and Length Exploitation, where sequence-averaged advantages reward either filler-token padding or premature truncation. Notably, distillation from Qwen3-4B-GRPO leaves AIME 2025 accuracy near 2%, showing that teacher capability alone does not determine teaching quality. The authors regulate token advantages with hard clipping, ãt=clip(∆ℓt,cmin,cmax), or soft log-scale compression, ãt=sign(∆ℓt)·log(1+|∆ℓt|), requiring no additional off-policy compute. With Qwen3-1.7B-GRPO as teacher, log-scale compression raises the student’s overall average from 28.1 for naive OPD to 30.4, while regulated OPD can outperform methods using a massive 30B teacher, establishing signal fidelity and capacity matching—not brute-force teacher scale—as the central determinants of successful OPD.

Original abstract

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis