NTH

On-Policy Self-Distillation without Any Supervision

AuthorsYijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos

August 15, 2026 2 min read
Watch on YouTube
The one-line take

U-OPSD lets an LLM improve its mathematical reasoning by learning from the agreement and mistakes in its own sampled solutions, without external labels or teacher models.

Key results

30k
OpenThoughts training subset

Unlabeled problem statements used for U-OPSD training

8
Rollouts per prompt

Independent on-policy generations sampled for majority voting

0.5
Self-consistency threshold

Minimum vote fraction required to form a pseudo-solution

8.5%
Qwen3-4B non-thinking improvement

Average gain over the base model across five math benchmarks

10.7%
Qwen3-8B non-thinking improvement

Average gain over the base model across five math benchmarks

13.3%
Incorrect pseudo-label rate

Wrong pseudo-labels measured in the in-domain probe

What the paper found

U-OPSD, or unsupervised on-policy self-distillation, removes the ground-truth solutions normally required by methods such as OPSD and GRPO. For each unlabeled problem, the model samples 8 Qwen3 rollouts, extracts their final answers, and accepts a pseudo-solution when the majority reaches a self-consistency threshold of 0.5. The longest agreeing reasoning trace becomes a detached teacher context, while the model applies full-vocabulary forward-KL distillation to disagreeing rollouts, targeting the prefixes where it contradicts its own consensus. Training uses only problem statements from a 30k subset of OpenThoughts, never its solutions. Across AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD raises average performance over Qwen3-4B and Qwen3-8B non-thinking baselines by 8.5% and 10.7%, respectively, and reaches average scores of 49.49 and 54.31, exceeding supervised OPSD by 3.2% and 2.3%. In thinking mode, gains shrink to 2.2% and 1.9% because the base models leave less headroom, but U-OPSD remains comparable to OPSD and surpasses GRPO. The method also transfers to Qwen3-30B-A3B-Instruct-2507 and Qwen3-4B-Instruct-2507, suggesting scalability beyond a single model size. Its central limitation is error inheritance: 13.3% of pseudo-labels were wrong in an in-domain probe. The result is relevant to developers at companies such as ByteDance and to the broader ecosystem around models like ChatGPT, Claude, Gemini, and DeepSeek because it offers a label-free alternative to costly supervised post-training.

Original abstract

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We propose Unsupervised On-Policy Self-Distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseudo-solution by majority vote under a self-consistency threshold. It then conditions a teacher distribution on the shortest pseudo-solution and distills it into prefixes of the model's longest incorrect completion, allowing the model to correct itself precisely where it is confidently wrong. Across diverse benchmarks, base models, and training settings, U-OPSD consistently improves over the base models and matches or surpasses supervised methods with ground truth (GT), such as OPSD and GRPO. On AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD improves over the base model by 8.5% and 10.7% on Qwen3 non-thinking mode at the 4B and 8B scales, respectively, and outperforms OPSD by an average of 3.2% and 2.3%. In thinking mode, U-OPSD remains on par with OPSD, outperforming it by 0.9% at 4B and matching it at 8B, while surpassing GRPO by 0.7% and 1.1%, respectively.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis