On-Policy Self-Distillation without Any Supervision
AuthorsYijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos
Resources
U-OPSD lets an LLM improve its mathematical reasoning by learning from the agreement and mistakes in its own sampled solutions, without external labels or teacher models.
Key results
Unlabeled problem statements used for U-OPSD training
Independent on-policy generations sampled for majority voting
Minimum vote fraction required to form a pseudo-solution
Average gain over the base model across five math benchmarks
Average gain over the base model across five math benchmarks
Wrong pseudo-labels measured in the in-domain probe
What the paper found
U-OPSD, or unsupervised on-policy self-distillation, removes the ground-truth solutions normally required by methods such as OPSD and GRPO. For each unlabeled problem, the model samples 8 Qwen3 rollouts, extracts their final answers, and accepts a pseudo-solution when the majority reaches a self-consistency threshold of 0.5. The longest agreeing reasoning trace becomes a detached teacher context, while the model applies full-vocabulary forward-KL distillation to disagreeing rollouts, targeting the prefixes where it contradicts its own consensus. Training uses only problem statements from a 30k subset of OpenThoughts, never its solutions. Across AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD raises average performance over Qwen3-4B and Qwen3-8B non-thinking baselines by 8.5% and 10.7%, respectively, and reaches average scores of 49.49 and 54.31, exceeding supervised OPSD by 3.2% and 2.3%. In thinking mode, gains shrink to 2.2% and 1.9% because the base models leave less headroom, but U-OPSD remains comparable to OPSD and surpasses GRPO. The method also transfers to Qwen3-30B-A3B-Instruct-2507 and Qwen3-4B-Instruct-2507, suggesting scalability beyond a single model size. Its central limitation is error inheritance: 13.3% of pseudo-labels were wrong in an in-domain probe. The result is relevant to developers at companies such as ByteDance and to the broader ecosystem around models like ChatGPT, Claude, Gemini, and DeepSeek because it offers a label-free alternative to costly supervised post-training.
Original abstract
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We propose Unsupervised On-Policy Self-Distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseudo-solution by majority vote under a self-consistency threshold. It then conditions a teacher distribution on the shortest pseudo-solution and distills it into prefixes of the model's longest incorrect completion, allowing the model to correct itself precisely where it is confidently wrong. Across diverse benchmarks, base models, and training settings, U-OPSD consistently improves over the base models and matches or surpasses supervised methods with ground truth (GT), such as OPSD and GRPO. On AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD improves over the base model by 8.5% and 10.7% on Qwen3 non-thinking mode at the 4B and 8B scales, respectively, and outperforms OPSD by an average of 3.2% and 2.3%. In thinking mode, U-OPSD remains on par with OPSD, outperforming it by 0.9% at 4B and matching it at 8B, while surpassing GRPO by 0.7% and 1.1%, respectively.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.