On-Policy Distillation with Negative-Policy Rollouts
AuthorsJaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo
AffiliationsNAVER AI Lab
NP-OPD improves language-model distillation by exposing students to rollouts from a weaker policy, helping them learn from the teacher while moving away from inferior behavior.
Key results
Prompts used for training across math, science, and code.
Non-thinking score, compared with 48.3 for standard OPD.
Non-thinking score, compared with 40.2 for standard OPD.
Thinking-mode score, compared with 63.0 for standard OPD.
Qwen3-1.7B speedup using pre-generated negative-policy rollouts, excluding their generation.
What the paper found
On-policy distillation usually trains a student to follow a stronger teacher, but can offer weak guidance where their output distributions differ. Negative-Policy OPD addresses this by mixing in rollouts from a lower-capability model, then applying the original teacher-student token-level reward to those trajectories—adding a move-away signal without changing the reward itself. Across 13 math, code, and science benchmarks, experiments with Qwen3 show that this approach improves standard OPD and also complements ExOPD and OPD2; tests on Gemma-4 extend the results beyond the Qwen3 family. For Qwen3-1.7B in non-thinking mode, the math average rises from 48.3 to 56.3, while Qwen3-4B’s code average increases from 40.2 to 51.1. Gemma-4-E4B-it’s math average improves from 63.0 to 66.4. The training set contains 30K prompts, and analyses find that the method concentrates probability suppression on tokens favored by the negative policy over the teacher. Because those rollouts can be pre-generated and reused, Qwen3-1.7B training on NVIDIA H100s runs 2.61× faster in thinking mode when using only negative-policy rollouts, excluding their initial generation. The central finding is that changing which trajectories receive teacher supervision can supply a useful negative reference while preserving OPD’s original learning signal.
Original abstract
On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.
Read the original paperMore in Large Language Models
Browse all 84 papers →HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang
OnePO uses temporary teacher guidance and reinforcement learning alone to turn general LLMs into stronger medical specialists without the usual supervised fine-tuning stage.
Learning to Learn a Language
Lennart Carstens-Behrens, Holger Fröhlich
A transformer trained on synthetic worlds learns to infer the hidden rules of real language and other sequential data without ever seeing language during training.
Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.