Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
AuthorsYi Ding, Ruqi Zhang
Resources
The paper argues that on-policy distillation may improve reasoning less by copying a teacher than by suppressing unlikely tokens, enabling a simpler teacher-free alternative that performs even better.
Key results
Overall teacher-supervision noise rate on Qwen3-1.7B student trajectories
Lowest student log-probability tokens receiving updates
Improved from the base score of 13.44
Relative improvement over the Qwen3-1.7B base model
Compared with 40.00 for the base model
Avg@32 point advantage over OPD
What the paper found
This paper challenges the standard interpretation of on-policy distillation, or OPD, as knowledge transfer from a stronger teacher. Using Qwen3-1.7B students, Qwen3-4B, Qwen3-30B-A3B, and Qwen3-235B-A22B teachers, the authors find that teacher advantages on student-generated trajectories are highly noisy: the overall noise rate rises from 30.6% with the 4B teacher to 50.6% with the 235B-A22B teacher. Yet training on noisy trajectories alone performs comparably to standard OPD, suggesting that improvement does not primarily come from matching teacher behavior. Gradient and ablation analyses show that learning concentrates on the lowest-log-probability tokens; replacing teacher advantages with a fixed negative advantage of -0.5 produces similar gains, while positive advantages can trigger collapse. The proposed On-Policy Self-Adaptation, or OPSA, removes teachers, labels, verifiable rewards, and reference answers. It updates only the lowest 20% of student log-probability tokens and scales negative advantages with token entropy, suppressing unlikely tail tokens while redistributing probability among plausible head tokens at high-entropy reasoning forks. On Qwen3-1.7B trained with DAPO-17k, OPSA raises AIME24 Avg@32 from 13.44 to 48.85, a 263.5% relative gain, and reaches 80.00 Pass@32 versus 40.00 for the base model. It also exceeds OPD by 16.77 points in AIME24 Avg@32, while improving AIME25, HMMT25, MBPP+, and GPQA-Diamond. The results imply that OPD’s K1-estimator gains may largely reflect self-directed probability reshaping rather than distillation.
Original abstract
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.