NTH

On-Policy Distillation with Negative-Policy Rollouts

AuthorsJaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

AffiliationsNAVER AI Lab

October 10, 2026 2 min read
Watch on YouTube
The one-line take

NP-OPD improves language-model distillation by exposing students to rollouts from a weaker policy, helping them learn from the teacher while moving away from inferior behavior.

Key results

30K
Training prompts

Prompts used for training across math, science, and code.

56.3
Qwen3-1.7B Math average with NP-OPD

Non-thinking score, compared with 48.3 for standard OPD.

51.1
Qwen3-4B Code average with NP-OPD

Non-thinking score, compared with 40.2 for standard OPD.

66.4
Gemma-4-E4B-it Math average with NP-OPD

Thinking-mode score, compared with 63.0 for standard OPD.

2.61
Thinking-mode training speedup

Qwen3-1.7B speedup using pre-generated negative-policy rollouts, excluding their generation.

What the paper found

On-policy distillation usually trains a student to follow a stronger teacher, but can offer weak guidance where their output distributions differ. Negative-Policy OPD addresses this by mixing in rollouts from a lower-capability model, then applying the original teacher-student token-level reward to those trajectories—adding a move-away signal without changing the reward itself. Across 13 math, code, and science benchmarks, experiments with Qwen3 show that this approach improves standard OPD and also complements ExOPD and OPD2; tests on Gemma-4 extend the results beyond the Qwen3 family. For Qwen3-1.7B in non-thinking mode, the math average rises from 48.3 to 56.3, while Qwen3-4B’s code average increases from 40.2 to 51.1. Gemma-4-E4B-it’s math average improves from 63.0 to 66.4. The training set contains 30K prompts, and analyses find that the method concentrates probability suppression on tokens favored by the negative policy over the teacher. Because those rollouts can be pre-generated and reused, Qwen3-1.7B training on NVIDIA H100s runs 2.61× faster in thinking mode when using only negative-policy rollouts, excluding their initial generation. The central finding is that changing which trajectories receive teacher supervision can supply a useful negative reference while preserving OPD’s original learning signal.

Original abstract

On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.

Read the original paper

More in Large Language Models

Browse all 84 papers →
01Llm

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang

OnePO uses temporary teacher guidance and reinforcement learning alone to turn general LLMs into stronger medical specialists without the usual supervised fine-tuning stage.

Read analysis
02Llm

Learning to Learn a Language

Lennart Carstens-Behrens, Holger Fröhlich

A transformer trained on synthetic worlds learns to infer the hidden rules of real language and other sequential data without ever seeing language during training.

Read analysis