NTH

Pruning and Distilling Mixture-of-Experts into Dense Language Models

AuthorsJunhyuck Kim, Jihun Yun, Haechan Kim, Gyeongman Kim, Joonghyun Bae, Jaewoong Cho

May 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to turn large mixture-of-experts language models into smaller dense models that keep much of the performance while being easier and cheaper to deploy.

Key results

350
Qwen3-30B-A3B configs

The main sweep evaluates 7 scoring methods, 5 grouping methods, 2 down-projection scalings, and K∈{8,16,32,64,128}, yielding 350 configurations on Qwen3-30B-A3B.

43.41%
Best 0.3B-token accuracy

DO-ACP with K=8 achieves the highest average downstream accuracy after 0.3B-token distillation on Qwen3-30B-A3B.

33.28%
Dense-to-dense baseline accuracy

The matched-parameter dense-to-dense pruning baseline reaches 33.28% average downstream accuracy after 0.3B-token distillation.

6.3%
Accuracy gain over dense pruning

MoE-to-dense outperforms dense-to-dense pruning by +6.3 percentage points after about 4B tokens of distillation.

1.6x
Training speedup

Under matched parameter count, MoE-to-dense trains 1.6× faster than dense-to-dense pruning, measured as 73 s/step versus 116 s/step.

58.10%
Extended-training accuracy

After about 4B tokens, DO-ACP with K=8 reaches 58.10% average downstream accuracy on Qwen3-30B-A3B.

What the paper found

This paper introduces the first systematic MoE-to-dense compression pipeline for language models, converting a trained Mixture-of-Experts teacher into a standard dense FFN through expert scoring, top-K selection, optional grouping/merging, block concatenation, and forward-KL distillation. The core novelty is DO-ACP, a D-optimal expert selector that maximizes the log-determinant of an importance-weighted Gram matrix, combining activation-weighted conditional probability with explicit redundancy penalty; greedy selection provides a 1−1/e approximation guarantee. On Qwen3-30B-A3B, the authors sweep 350 configurations across 7 scoring methods, 5 grouping methods, 2 down-projection scalings, and K∈{8,16,32,64,128}. Scoring dominates: DO-ACP reaches 43.41% average accuracy after 0.3B tokens of distillation, versus 37.17% for reverse-KL-style weak scoring baselines and 33.28% for a strong dense-to-dense pruning baseline from Qwen3-32B, a +10.1 point margin. Under matched parameter count, MoE-to-dense also trains 1.6× faster than dense-to-dense pruning, 73 s/step versus 116 s/step. After ∼4B tokens, DO-ACP scales to 58.10% average accuracy, beating dense-to-dense by +6.3 points and random FFN initialization by +12.7 points. The method generalizes to DeepSeek-V2-Lite and GPT-OSS-20B, where DO-ACP remains best and pure pruning at K=k is consistently optimal, showing that expert diversity, not averaging, is the key signal for converting MoE capacity into a compact dense student.

Original abstract

Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment. Existing compression methods reduce the number of experts but the output remains an MoE model with the same fundamental limitation. We present the first systematic framework for converting a trained MoE into a standard fully dense architecture: experts are scored, selected, and grouped, then concatenated into a dense FFN and refined by knowledge distillation from the MoE teacher. We evaluate 7 scoring, 5 grouping, and 2 magnitude scaling methods across a range of selected expert counts on Qwen3-30B-A3B, yielding 350 configurations. We find that the choice of scoring method is the most impactful, with our novel diversity-aware scoring consistently outperforming prior methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and GPT-OSS-20B. Under a controlled comparison at matched parameter count, MoE-to-dense outperforms dense-to-dense pruning by +6.3 pp in average downstream accuracy after ~4B-token distillation at 1.6x faster training wall-clock speed.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis