Pruning and Distilling Mixture-of-Experts into Dense Language Models
AuthorsJunhyuck Kim, Jihun Yun, Haechan Kim, Gyeongman Kim, Joonghyun Bae, Jaewoong Cho
Resources
This paper shows how to turn large mixture-of-experts language models into smaller dense models that keep much of the performance while being easier and cheaper to deploy.
Key results
The main sweep evaluates 7 scoring methods, 5 grouping methods, 2 down-projection scalings, and K∈{8,16,32,64,128}, yielding 350 configurations on Qwen3-30B-A3B.
DO-ACP with K=8 achieves the highest average downstream accuracy after 0.3B-token distillation on Qwen3-30B-A3B.
The matched-parameter dense-to-dense pruning baseline reaches 33.28% average downstream accuracy after 0.3B-token distillation.
MoE-to-dense outperforms dense-to-dense pruning by +6.3 percentage points after about 4B tokens of distillation.
Under matched parameter count, MoE-to-dense trains 1.6× faster than dense-to-dense pruning, measured as 73 s/step versus 116 s/step.
After about 4B tokens, DO-ACP with K=8 reaches 58.10% average downstream accuracy on Qwen3-30B-A3B.
What the paper found
This paper introduces the first systematic MoE-to-dense compression pipeline for language models, converting a trained Mixture-of-Experts teacher into a standard dense FFN through expert scoring, top-K selection, optional grouping/merging, block concatenation, and forward-KL distillation. The core novelty is DO-ACP, a D-optimal expert selector that maximizes the log-determinant of an importance-weighted Gram matrix, combining activation-weighted conditional probability with explicit redundancy penalty; greedy selection provides a 1−1/e approximation guarantee. On Qwen3-30B-A3B, the authors sweep 350 configurations across 7 scoring methods, 5 grouping methods, 2 down-projection scalings, and K∈{8,16,32,64,128}. Scoring dominates: DO-ACP reaches 43.41% average accuracy after 0.3B tokens of distillation, versus 37.17% for reverse-KL-style weak scoring baselines and 33.28% for a strong dense-to-dense pruning baseline from Qwen3-32B, a +10.1 point margin. Under matched parameter count, MoE-to-dense also trains 1.6× faster than dense-to-dense pruning, 73 s/step versus 116 s/step. After ∼4B tokens, DO-ACP scales to 58.10% average accuracy, beating dense-to-dense by +6.3 points and random FFN initialization by +12.7 points. The method generalizes to DeepSeek-V2-Lite and GPT-OSS-20B, where DO-ACP remains best and pure pruning at K=k is consistently optimal, showing that expert diversity, not averaging, is the key signal for converting MoE capacity into a compact dense student.
Original abstract
Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment. Existing compression methods reduce the number of experts but the output remains an MoE model with the same fundamental limitation. We present the first systematic framework for converting a trained MoE into a standard fully dense architecture: experts are scored, selected, and grouped, then concatenated into a dense FFN and refined by knowledge distillation from the MoE teacher. We evaluate 7 scoring, 5 grouping, and 2 magnitude scaling methods across a range of selected expert counts on Qwen3-30B-A3B, yielding 350 configurations. We find that the choice of scoring method is the most impactful, with our novel diversity-aware scoring consistently outperforming prior methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and GPT-OSS-20B. Under a controlled comparison at matched parameter count, MoE-to-dense outperforms dense-to-dense pruning by +6.3 pp in average downstream accuracy after ~4B-token distillation at 1.6x faster training wall-clock speed.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.