NTH

MobileMoE: Scaling On-Device Mixture of Experts

AuthorsYanbei Chen, Hanxian Huang, Ernie Chang, Jacob Szwejbka, Digant Desai, Zechun Liu, Vikas Chandra, Raghuraman Krishnamoorthi

June 9, 2026 2 min read
Watch on YouTube
The one-line take

MobileMoE shows that sparse mixture-of-experts language models can run efficiently on phones, delivering better speed-memory tradeoffs than dense on-device LLMs.

Key results

0.3B
MobileMoE-S active params

small on-device model scale from the derived family

5.3B
MobileMoE-L total params

largest MobileMoE model size after architecture search

6T
Pre-training tokens

token budget used for MobileMoE pre-training

500B
Mid-training tokens

context-length extension and domain sharpening stage

60.1
MobileMoE-L overall accuracy

average score after instruction fine-tuning on 14 benchmarks

3.8x
MobileMoE-S prefill speedup

fastest reported prefill gain over MobileLLM-Pro on iPhone 16 Pro

What the paper found

MobileMoE is a Meta AI on-device Mixture-of-Experts language-model family that attacks the sub-billion active-parameter regime with a mobile-specific scaling law rather than a server-scale recipe. The paper derives a generalized loss model under joint memory and compute constraints, then uses it to select a sweet spot of moderate sparsity, fine-grained experts, and one shared expert, yielding MobileMoE-S/M/L at 0.3B/0.5B/0.9B active parameters and 1.3B/2.8B/5.3B total parameters. Trained in four stages on open data, including about 6T pre-training tokens and 500B mid-training tokens, the models achieve a new Pareto frontier across 14 benchmarks: MobileMoE-L reaches 59.8 average accuracy after mid-training and 60.1 after instruction tuning, while matching or surpassing OLMoE-1B-7B with 30% fewer active parameters and 23% smaller total size. The quantized INT4 versions preserve most quality, with MobileMoE-L still at 57.8 average accuracy and only 2.75 GB weight memory. To prove real device viability, Meta also built a fused MoE kernel in ExecuTorch and profiled Samsung Galaxy S25 and iPhone 16 Pro; MobileMoE-S runs 1.8–3.8× faster in prefill and 2.2–3.4× faster in decode than MobileLLM-Pro, while staying within smartphone DRAM budgets and showing that sparse expert routing can beat dense mobile LLMs on both accuracy and latency.

Original abstract

Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot - moderate sparsity with fine-grained and shared experts - that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4$\times$ fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable INT4 weight memory, MobileMoE-S delivers $1.8$-$3.8\times$ faster prefill and $2.2$-$3.4\times$ faster decode than the dense baseline MobileLLM-Pro.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis