NTH

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

AuthorsZhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li

AffiliationsCase Western Reserve University Zillow Group, Inc. University of Illinois at Chicago · University of Florida

October 8, 2026 2 min read
Watch on YouTube
The one-line take

Latent-MOPD teaches one language model to combine several specialists by learning both what they predict and how they represent it.

Key results

1.5B
Same-family student size

DeepSeek-R1-Distill-Qwen student distilled from same-family specialists

17,856
Training prompt pool

Prompts drawn from math, code, and logic domains

1.05
Normalized score

Latent-MOPD score in the same-family setting

39.4
AIME24 cross-family score

Score with 7B Qwen-based teachers, versus 36.5 for token-only MOPD

62
Distillation steps

Optimizer-step budget used in the main experiments

What the paper found

Latent-MOPD introduces representation-level multi-teacher on-policy distillation for language models, extending MOPD beyond token probabilities by matching the hidden states that produce those predictions. A domain router sends each student-generated response to a math, code, or logic specialist, and the same teacher supplies both token-level policy-gradient targets and late-layer representation targets. For same-family models, the method aligns the final three layers directly; for cross-family teachers, it uses a shared trainable projection to bridge hidden widths. Domain-pure optimizer updates prevent conflicting representation gradients, while a per-teacher crossfade gradually replaces latent supervision with token supervision. Experiments distill three specialists into the 1.5B DeepSeek-R1-Distill-Qwen student, using 17,856 prompts from DAPO-Math-17k, OpenCodeReasoning, and Reasoning Gym. Across nine math, coding, and logic benchmarks, Latent-MOPD beats token-only, representation-only, and uniform-averaging baselines, reaches a normalized score of 1.05, and surpasses the strongest individual teacher on five benchmarks despite matching each teacher’s parameter count. With 7B Qwen-based teachers, including NVIDIA’s AceReason-Nemotron and DeepSeek’s R1-Distill-Qwen, the method also improves on all six benchmarks; AIME24 rises to 39.4 from 36.5 for token-only MOPD. Training lasts 62 optimizer steps, and the merged-initialization experiment improves five of six benchmarks, suggesting that latent supervision can consolidate specialist capabilities into one deployable student without retaining teachers or projection layers at inference.

Original abstract

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis