Rethinking Self-Distillation for Multi-Teacher Capability Merging
AuthorsRoy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
Affiliations[
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.
Key results
MOPD’s full-run training GPU-hours divided by SFT’s on MIA.
MOPD’s full-run training GPU-hours divided by SFT’s on Agent×3.
Teacher capability gains recovered on average across the two task compositions.
Teacher capability gains recovered on average across the two task compositions.
What the paper found
As multi-teacher distillation appears in systems such as DeepSeek-V4 and GLM-5, this study tests whether on-policy training is actually needed to combine expert capabilities. Comparing supervised fine-tuning, soft-label distillation, a hybrid offline method, and multi-teacher on-policy distillation across four Qwen and Ministral models and 11 benchmarks, the researchers find that independently tuned methods achieve nearly identical accuracy. Yet MOPD requires 14.8 times SFT’s training GPU-hours on MIA and 23.1 times on Agent×3, where student rollouts also require live environment interaction. The authors argue that some published on-policy advantages shrink when SFT uses successful, rejection-sampled teacher trajectories and gets its own learning-rate tuning; this contrasts with approaches such as Microsoft’s MAI-Thinking-1, which uses SFT on teacher-generated data. As a training-free alternative, weight merging recovers 99.0% of teacher gains with Qwen3-8B, but only 87.3% with Qwen3-1.7B, and works less reliably when expert tasks overlap. The central takeaway is that careful off-policy tuning—or merging when models and tasks are compatible—can provide capability integration without MOPD’s substantial rollout cost.
Original abstract
Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
LSPD uses reinforcement-learning ideas to make LLM policy distillation more sample-efficient while preserving the diversity needed for stronger reasoning.