MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
AuthorsWenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, Fuli Luo
Resources
MOPD is a new way to merge multiple specialized LLM skills by distilling several RL-trained teachers into one student using the student’s own rollouts, and it appears to work well at frontier scale.
Key results
MOPD final aggregate integration score on the three-domain benchmark
Best baseline on Qwen3-30B-A3B
Sequential domain-training baseline on Qwen3-30B-A3B
Offline teacher-rollout fine-tuning baseline on Qwen3-30B-A3B
Best Param-Merge variant on Qwen3-30B-A3B
Approximate instruction-following sample count to reach teacher-level plateau
What the paper found
MOPD, or Multi-Teacher On-Policy Distillation, is a post-training method from Peking University and Xiaomi’s LLM Core that integrates multiple specialized RL capabilities into one LLM without weight-space merging or off-policy imitation. The pipeline is three-stage: a shared SFT checkpoint is trained first, then separate domain teachers are learned in parallel with domain-specific RL, and finally a student is distilled on its own rollouts using token-level reverse KL against the routed teacher for each prompt. On Qwen3-30B-A3B, evaluated across math, instruction following, and software engineering, MOPD achieves a normalized score of 0.9373, beating Mix-RL at 0.8818, Cascade RL at 0.7752, Off-Policy Finetune at 0.8241, and Param-Merge task arithmetic at 0.8574, while closing 91% to 95% of the student-teacher headroom on each domain. The paper also shows a sample-efficiency gain: MOPD reaches teacher-level plateaus on instruction following with about 25K samples and software engineering with about 30K, versus roughly 150K to 180K for Mix-RL. On the industrial-scale MiMo-V2-Flash model, MOPD matches or exceeds the corresponding teacher on most benchmarks, including AIME25 at 94.1, HMMT25 at 84.4, and 𝜏2-Telecom at 95.3. Analysis finds that policy-gradient and top-k distillation are comparable when teacher and student share the same origin, but replacing the teacher with the distributionally different Qwen3-235B-A22B causes instability and even catastrophic collapse, underscoring that same-origin teachers are critical for stable integration.
Original abstract
Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose performance. In this work, we propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm for combining the capabilities of multiple domain RL teachers: we first run per-domain specialised RL to obtain a set of domain teachers, then distill these teachers into the student on its own rollouts. This eliminates exposure bias and provides a dense optimization signal. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher's capability. MOPD also enables parallel, independent development of domain teachers, removing the cross-domain coupling typical of multi-domain post-training. MOPD has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model, demonstrating its practical value for capability integration in frontier-scale LLMs.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.