NTH

Scaling Laws for Looped Mixture of Experts

AuthorsYanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

AffiliationsMeta AI

October 9, 2026 2 min read
Watch on YouTube
The one-line take

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Key results

14
Downstream benchmarks

Benchmarks used to evaluate the complementary effects of sparsity and recurrence.

3
Active-parameter efficiency

Sparsity achieved this fold efficiency on overall benchmark performance.

2
Reasoning parameter efficiency

Recurrence achieved this fold efficiency in total parameters on reasoning.

10.3
Inference-time overall score gain

Score-point increase when recurrence varied from R=1 to R=5.

What the paper found

Meta’s study introduces Loop Scaling Laws, a framework that jointly models model size, training data, recurrence, and Mixture-of-Experts sparsity. Its key innovation is a bounded effective-parameter mapping: recurrent passes add diminishing capacity, while greater sparsity raises and prolongs that gain. This captures how expert routing lets tokens reach different parameters on successive passes, and it predicts held-out loss better than unbounded linear or power-law alternatives. The law can also select recurrence and expert counts under compute and memory limits. Across 14 benchmarks, sparsity delivered 3-fold active-parameter efficiency on overall performance, while recurrence delivered 2-fold total-parameter efficiency on reasoning. At matched training compute, a 0.3B-active, 1.3B-total looped MoE matched the reasoning performance of a roughly twice-larger 0.6B-active, 2.9B-total MoE; varying recurrence at inference raised overall benchmark scores by 10.3 points, from 36.6 to 46.9. The approach offers a way to co-design capacity and depth in the same broad MoE landscape as DeepSeek-V3, Mixtral, Qwen3MoE, and Gemini; those models were not evaluated in this study.

Original abstract

Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
03Efficiency

Disaggregated Quantization: Specializing LLM Prefill and Decode

Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

Disaggregated quantization gives LLM prefill and decode their own specialized weights and formats, improving low-bit accuracy while speeding up first-token generation.

Read analysis