Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
AuthorsZhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li
AffiliationsCase Western Reserve University Zillow Group, Inc. University of Illinois at Chicago · University of Florida
Resources
Latent-MOPD teaches one language model to combine several specialists by learning both what they predict and how they represent it.
Key results
DeepSeek-R1-Distill-Qwen student distilled from same-family specialists
Prompts drawn from math, code, and logic domains
Latent-MOPD score in the same-family setting
Score with 7B Qwen-based teachers, versus 36.5 for token-only MOPD
Optimizer-step budget used in the main experiments
What the paper found
Latent-MOPD introduces representation-level multi-teacher on-policy distillation for language models, extending MOPD beyond token probabilities by matching the hidden states that produce those predictions. A domain router sends each student-generated response to a math, code, or logic specialist, and the same teacher supplies both token-level policy-gradient targets and late-layer representation targets. For same-family models, the method aligns the final three layers directly; for cross-family teachers, it uses a shared trainable projection to bridge hidden widths. Domain-pure optimizer updates prevent conflicting representation gradients, while a per-teacher crossfade gradually replaces latent supervision with token supervision. Experiments distill three specialists into the 1.5B DeepSeek-R1-Distill-Qwen student, using 17,856 prompts from DAPO-Math-17k, OpenCodeReasoning, and Reasoning Gym. Across nine math, coding, and logic benchmarks, Latent-MOPD beats token-only, representation-only, and uniform-averaging baselines, reaches a normalized score of 1.05, and surpasses the strongest individual teacher on five benchmarks despite matching each teacher’s parameter count. With 7B Qwen-based teachers, including NVIDIA’s AceReason-Nemotron and DeepSeek’s R1-Distill-Qwen, the method also improves on all six benchmarks; AIME24 rises to 39.4 from 36.5 for token-only MOPD. Training lasts 62 optimizer steps, and the merged-initialization experiment improves five of six benchmarks, suggesting that latent supervision can consolidate specialist capabilities into one deployable student without retaining teachers or projection layers at inference.
Original abstract
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.