NTH

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

AuthorsYuhang Zhou, Lizhu Zhang, Yifan Wu, Mingyi Wang, Peng Bo, Jiayi Liu, Xiangjun Fan, Zhuokai Zhao

June 16, 2026 2 min read
Watch on YouTube
The one-line take

OmniOPD is a new way to train smaller models from stronger teachers without needing token logits, using chunk-level semantic verification and uncertainty-based checks to improve distillation, especially on math tasks.

Key results

45.31%
Math relative gain over SFT

Maximum relative improvement reported on mathematical reasoning versus offline SFT.

18.52%
Code relative gain over SFT

Maximum relative improvement reported on competitive programming versus offline SFT.

28.64%
Math gain over white-box OPD

Maximum relative improvement over standard white-box OPD on mathematical reasoning.

69.08%
Qwen3-4B OmniOPD average math

Average Pass@1 for Qwen3-4B distilled from Qwen3-32B on math benchmarks.

63.80%
Qwen3-4B SFT average math

Offline SFT baseline for the Qwen3-4B student on math benchmarks.

8.28%
KL ablation average math

Average Pass@1 after removing the KL anchor, showing catastrophic collapse.

What the paper found

OmniOPD from Meta AI introduces a logit-free on-policy distillation method that replaces brittle token-level reverse-KL matching with speculative verification over semantic chunks, making it usable with proprietary teachers such as Claude-4.5-Haiku and Gemini-2.5-Flash that expose only text. The student generates its own reasoning trajectory, OmniOPD audits only the highest-entropy decision forks, then queries the teacher for Monte Carlo rollouts over 50-token chunks and scores them with a semantic similarity metric such as Edit Distance or ROUGE-1; a Dirichlet-Multinomial Bayesian prior and a base-model KL anchor stabilize the sparse signal. On Qwen3 students, the method improves mathematical reasoning by up to 45.31% relative over SFT and competitive programming by up to 18.52% relative, while beating standard white-box OPD by up to 28.64% on math. In the main Qwen3-4B to Qwen3-32B setup, OmniOPD reaches 69.08% average Pass@1 on math versus 63.80% for SFT and 64.16% for OPD, and the strongest black-box teacher setting with Gemini-2.5-Flash reaches 75.67%. Ablations show the KL anchor is essential: removing it collapses average math accuracy from 69.08% to 8.28%.

Original abstract

On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitigating both the off-policy distribution shift of Supervised Fine-Tuning (SFT) and the sparse credit assignment of Reinforcement Learning (RL). However, standard OPD faces two coupled limitations. First, it requires direct access to the teacher's token-level logits, excluding a broad class of capable proprietary models from serving as teachers. Second, the token-level logit signal itself is brittle, depending on a narrow overlap of plausible next tokens between teacher and student, and prone to amplifying degenerate patterns such as repetition loops. In this paper, we introduce OmniOPD, a novel framework that addresses both limitations through a logit-free, chunk-level supervision signal. OmniOPD replaces deterministic logit matching with Monte Carlo rollouts that approximate the teacher's local preferences through a continuous semantic similarity metric over multi-token chunks, and concentrates this supervision via a peak-entropy scheduler that audits the student only at its high-uncertainty reasoning forks. A Dirichlet-Multinomial Bayesian prior and a base-model KL anchor further bound the variance of discrete sampling and prevent policy collapse across unaudited tokens. Across competitive benchmarks, OmniOPD surpasses the standard OPD approach by up to +28.64% on math, confirming that chunk-level semantic verification extracts a more reliable learning signal than token-level logit matching, whose high information density is offset by significant noise and brittleness. Furthermore, when paired with stronger black-box teachers such as Claude-4.5-Haiku and Gemini-2.5-Flash, OmniOPD achieves an additional +9.54% relative on math over its open-weight teacher counterpart, advancing the student past the performance of self-exploratory RL.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis