NTH

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

AuthorsYongkang Yang, Zhezheng Hao, Hong Zhang, Yi Liu, Xiankun Lin, Wence Ji, Fanjunduo Wei, Jiarui Yu, Qiang Lin, Xiaoyun Liang, Hande Dong

August 15, 2026 2 min read
Watch on YouTube
The one-line take

USD teaches language models to choose both what supervision to absorb and how difficult that supervision should be, adapting self-distillation to the model’s changing learning capacity.

Key results

2.3
Avg@12 improvement

USD increases the nine-cell mathematical reasoning average from 56.4 to 58.7 over vanilla OPSD.

30K
Training data scale

Experiments use up to 30K OpenThoughts problem–solution pairs.

0.3
Optimal capacity budget

Qwen3-1.7B performance peaks at the learning-capacity budget ε=0.3.

1024
USD rollout length

USD uses a single 1024-token rollout per problem.

8
GRPO rollout count

The comparison uses 8 GRPO rollouts per problem, each up to 16K tokens.

What the paper found

This Tencent paper introduces Unified on-policy Self-Distillation, or USD, for improving reasoning in Qwen3 language models. In OPSD, the student learns from its own rollout while a same-model teacher receives privileged information, but two common improvements are usually optimized separately: token selection decides where supervision is applied, and privileged-information adaptation controls how much context the teacher sees. USD shows these choices are coupled by the student’s learning capacity. It maximizes teacher–student KL divergence under a budget on learning difficulty, using per-token divergence and teacher-choice surprise to compute soft weights; a single dual variable, lambda, simultaneously filters difficult tokens and adjusts the privileged-information strength through a primal-dual online update. The method adds only one O(T) bookkeeping pass, with no auxiliary network or extra forward pass. Across Qwen3-1.7B, 4B, and 8B, trained on up to 30K OpenThoughts problem–solution pairs and evaluated on AIME 2024, AIME 2025, and HMMT 2025, USD raises the nine-cell Avg@12 from 56.4 to 58.7, a 2.3-point gain over vanilla OPSD. The best capacity budget is ε=0.3, and the method uses a single 1024-token rollout per problem, compared with GRPO’s 8 rollouts of up to 16K tokens, while outperforming GRPO at every tested scale.

Original abstract

On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. However, we show that each line optimizes one variable while holding the other fixed, which leads to a suboptimal solution. We argue that the two variables are coupled through the student's learning capacity: the privileged information sets the per-token divergence the teacher prescribes, while token weighting selects which of these the student must absorb. We formalize the two lines of work into a unified optimization framework, which maximizes the aggregate teacher--student divergence, subject to a budget on the aggregate learning difficulty the student can absorb. Under this modelling, we propose Unified On-Policy Self-Distillation (USD), a lightweight online algorithm to solve the Lagrangian. USD reveals that a single dual variable governs both decisions: at one price for learning difficulty, it simultaneously sets the token-selection threshold and the direction of privileged-information adjustment, keeping supervision matched to the student's evolving capacity. Through extensive experiments, USD consistently demonstrates superior performance over OPSD and token- and PI-side baselines across various model scales on various reasoning benchmarks. Code is available at https://github.com/lauvlalala/USD.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis