NTH

Rethinking Self-Distillation for Multi-Teacher Capability Merging

AuthorsRoy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

Affiliations[

October 9, 2026 2 min read
Watch on YouTube
The one-line take

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Key results

14.8
MOPD-to-SFT GPU-hour ratio on MIA

MOPD’s full-run training GPU-hours divided by SFT’s on MIA.

23.1
MOPD-to-SFT GPU-hour ratio on Agent×3

MOPD’s full-run training GPU-hours divided by SFT’s on Agent×3.

99.0%
Qwen3-8B weight-merging recovery

Teacher capability gains recovered on average across the two task compositions.

87.3%
Qwen3-1.7B weight-merging recovery

Teacher capability gains recovered on average across the two task compositions.

What the paper found

As multi-teacher distillation appears in systems such as DeepSeek-V4 and GLM-5, this study tests whether on-policy training is actually needed to combine expert capabilities. Comparing supervised fine-tuning, soft-label distillation, a hybrid offline method, and multi-teacher on-policy distillation across four Qwen and Ministral models and 11 benchmarks, the researchers find that independently tuned methods achieve nearly identical accuracy. Yet MOPD requires 14.8 times SFT’s training GPU-hours on MIA and 23.1 times on Agent×3, where student rollouts also require live environment interaction. The authors argue that some published on-policy advantages shrink when SFT uses successful, rejection-sampled teacher trajectories and gets its own learning-rate tuning; this contrasts with approaches such as Microsoft’s MAI-Thinking-1, which uses SFT on teacher-generated data. As a training-free alternative, weight merging recovers 99.0% of teacher gains with Qwen3-8B, but only 87.3% with Qwen3-1.7B, and works less reliably when expert tasks overlap. The central takeaway is that careful off-policy tuning—or merging when models and tasks are compatible—can provide capability integration without MOPD’s substantial rollout cost.

Original abstract

Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis