NTH

Hyperparameter Scaling Laws Across MoE Sparsity

AuthorsChangxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

September 11, 2026 2 min read
Watch on YouTube
The one-line take

This work shows how to predict learning rates and batch sizes for increasingly sparse MoE models, making large-scale training more compute-efficient and transferable.

Key results

1800
Pre-training runs

Number of experiments used to fit and validate the MoE hyperparameter laws.

20T
Training tokens

Approximate token volume processed across the experimental sweep.

200000
Compute cost

Equivalent NVIDIA H800 GPU-hours used for the experiments.

0.1361
Learning-rate sparsity exponent

Fitted exponent of activation ratio A in the optimal learning-rate law.

-0.0841
Batch-size sparsity exponent

Fitted exponent of activation ratio A in the optimal batch-size law.

1/64
Held-out activation ratio

Expert activation ratio of the 12B-total-parameter extrapolation target.

What the paper found

This study shows that standard hyperparameter scaling rules break down for ultra-sparse Mixture-of-Experts models because sparsity changes the optimal learning rate and token batch size independently of activated or total parameter count. Across 1800 pre-training runs, the experiments processed 20T tokens at a cost of 200000 equivalent NVIDIA H800 GPU-hours, using Megatron, the Muon optimizer, and auxiliary-loss-free routing. At fixed sparsity, optimal learning rate scales with analytically computed training FLOPs C, while optimal batch size scales with training tokens D; activation ratio A then modifies each through a multiplicative power law: η*=0.8343 C^-0.1385 A^0.1361 and B*=6.4765 D^0.5181 A^-0.0841. The positive learning-rate sparsity exponent means denser activation supports a higher rate, while the negative batch-size exponent means greater sparsity requires larger global batches, consistent with increased per-expert gradient noise. The multiplicative form outperformed scale-only and additive alternatives in grouped validation and was compared with DeepSeek and Microsoft scaling laws. On a held-out 12B-total-parameter MoE activating only 1/64 of its experts, the predicted settings remained near the observed optimum, supporting joint extrapolation across model scale, training duration, compute, and sparsity. Additional controls showed transfer across expert granularities, indicating that activation ratio—not expert count alone—is the relevant predictor.

Original abstract

Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis