NTH

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

AuthorsNayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim

August 24, 2026 2 min read
Watch on YouTube
The one-line take

A compute-saving framework predicts the right learning rate for enormous MoE models by scaling up results from much smaller proxy runs.

Key results

0.95
Token-scaling regression R²

Log-log regression accurately extrapolated optimal learning rates across token horizons.

3.85e-4
Predicted 10T learning rate

Extrapolated optimal learning rate for 10T-token pretraining.

155B
Target total parameters

Total parameter count of the validated MoE foundation model.

17B
Target active parameters

Parameters activated per token in the validated MoE model.

10T
Training horizon

Tokens used for full-scale pretraining.

98
Proxy-search compute ratio

The target training required approximately 98 times the compute of the proxy runs.

What the paper found

Large-scale Mixture-of-Experts systems such as DeepSeek-V3, Qwen3-235B-A22B, Kimi-K2.5, and GLM-5 increase capacity efficiently, but learning-rate sweeps become impractical as model width, sparsity, and training tokens grow. This paper proposes a two-step transfer method: first, an MoE-specific Maximal Update Parameterization, or µP, combined with Multi-head Latent Attention and the Muon optimizer, transfers optimal learning rates across width and expert-count scaling while keeping depth, active experts, and expert dimensions fixed; second, a log-log token scaling law extrapolates the rate to long training horizons. Proxy runs use Warmup-Stable-Decay scheduling, Exponential Moving Average with smoothing factor 0.6, and quadratic validation-loss fits to estimate optimal rates at multiple token budgets. Regression achieves R² 0.95 and predicts a 3.85e-4 learning rate for 10T-token pretraining. The approach is validated on an MoE foundation model with 155B total and 17B active parameters, trained for 10T tokens on NVIDIA H200 GPUs, while requiring approximately 98 times less compute for the proxy-based search than exhaustive target-scale tuning. The resulting training trajectory remains stable, and evaluations including MMLU, MMLU-Pro, BBH, Global-MMLU, MATH, GSM8K, MBPP, and HumanEval place the model on a favorable compute-performance frontier against comparable open-weight MoEs.

Original abstract

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons. First, we formulate a Maximal Update Parameterization ($μ$P) adaptation for MoE architectures utilizing Multi-head Latent Attention (MLA) and the Muon optimizer, demonstrating that optimal learning rates transfer consistently across width-scaled models. Second, we extend this transferability along the token dimension by establishing a predictive scaling law. By applying linear regression to the optimal values derived from small proxy models on limited budgets, we successfully extrapolate the ideal learning rate to massive training horizons (e.g., 10 trillion tokens) with high fidelity ($R^2=0.95$). Consequently, this indicates that proxy training on small models is sufficient to determine the optimal learning rate for the extensive training of large-scale MoEs. We apply the proposed methodology to pretrain our foundation model (155B total, 17B active parameters) from scratch, and the stable training and evaluation results validate that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs.

Read the original paper

More in Optimization

Browse all 36 papers →