MobileMoE: Scaling On-Device Mixture of Experts
AuthorsYanbei Chen, Hanxian Huang, Ernie Chang, Jacob Szwejbka, Digant Desai, Zechun Liu, Vikas Chandra, Raghuraman Krishnamoorthi
Resources
MobileMoE shows that sparse mixture-of-experts language models can run efficiently on phones, delivering better speed-memory tradeoffs than dense on-device LLMs.
Key results
small on-device model scale from the derived family
largest MobileMoE model size after architecture search
token budget used for MobileMoE pre-training
context-length extension and domain sharpening stage
average score after instruction fine-tuning on 14 benchmarks
fastest reported prefill gain over MobileLLM-Pro on iPhone 16 Pro
What the paper found
MobileMoE is a Meta AI on-device Mixture-of-Experts language-model family that attacks the sub-billion active-parameter regime with a mobile-specific scaling law rather than a server-scale recipe. The paper derives a generalized loss model under joint memory and compute constraints, then uses it to select a sweet spot of moderate sparsity, fine-grained experts, and one shared expert, yielding MobileMoE-S/M/L at 0.3B/0.5B/0.9B active parameters and 1.3B/2.8B/5.3B total parameters. Trained in four stages on open data, including about 6T pre-training tokens and 500B mid-training tokens, the models achieve a new Pareto frontier across 14 benchmarks: MobileMoE-L reaches 59.8 average accuracy after mid-training and 60.1 after instruction tuning, while matching or surpassing OLMoE-1B-7B with 30% fewer active parameters and 23% smaller total size. The quantized INT4 versions preserve most quality, with MobileMoE-L still at 57.8 average accuracy and only 2.75 GB weight memory. To prove real device viability, Meta also built a fused MoE kernel in ExecuTorch and profiled Samsung Galaxy S25 and iPhone 16 Pro; MobileMoE-S runs 1.8–3.8× faster in prefill and 2.2–3.4× faster in decode than MobileLLM-Pro, while staying within smartphone DRAM budgets and showing that sparse expert routing can beat dense mobile LLMs on both accuracy and latency.
Original abstract
Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot - moderate sparsity with fine-grained and shared experts - that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4$\times$ fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable INT4 weight memory, MobileMoE-S delivers $1.8$-$3.8\times$ faster prefill and $2.2$-$3.4\times$ faster decode than the dense baseline MobileLLM-Pro.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.