NTH

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

AuthorsNuemaan Malik

July 25, 2026 3 min read
Watch on YouTube
The one-line take

SkewAdam cuts MoE optimizer memory dramatically by giving different parameter groups only the optimizer state they actually need, without sacrificing accuracy.

Key results

6.78B
Model size

Parameters in the evaluated MoE language model.

1.29GB
SkewAdam state

Optimizer state required by SkewAdam.

50.55GB
AdamW state

Optimizer state required by AdamW on the same model.

31.3GB
Peak memory

Peak training memory with SkewAdam, compared with 81.4GB for AdamW.

108.4
SkewAdam perplexity

Validation perplexity after 82M tokens.

82M
Training tokens

Approximate tokens processed in the controlled comparison.

What the paper found

Nuemaan Malik’s paper proposes SkewAdam, a tiered optimizer for Mixture-of-Experts training that assigns state according to parameter role rather than treating an MoE like a homogeneous model. In a 6.78B-parameter decoder trained on OpenWebText, the dense backbone keeps float32 momentum plus factored second moments, the 95% expert bank keeps only factored second moments, and the tiny router retains an exact second moment for load-balancing stability. This reduces optimizer state to 1.29GB versus AdamW’s 50.55GB and lowers peak memory from 81.4GB to 31.3GB, making the run fit within a 40GB accelerator while sustaining 5,000 tokens per second on an NVIDIA H200. After 82M tokens, SkewAdam reaches validation perplexity 108.4, outperforming AdamW at 126.8, Muon at 120.2, and Lion at 393.7; routing remains within 1% of the uniform balance floor. Ablations show that uniform momentum with factored variance matches SkewAdam’s accuracy at twenty times the state, while restoring momentum to experts adds 24GB with negligible perplexity benefit. The paper’s central finding is therefore allocation, not a fundamentally superior update rule: momentum matters for the dense backbone, but is largely wasted on sparsely activated experts. Comparisons use NVIDIA H200 and H100 systems, with results positioned alongside MoE systems such as Mixtral and DeepSeekMoE rather than claiming production-scale validation.

Original abstract

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the three parameter populations of an MoE - the dense backbone, the experts, and the router - differ enough in size and gradient statistics that they should not receive the same state. SkewAdam keeps float32 momentum plus a factored second moment for the backbone (5% of parameters), a factored second moment alone for the experts (95%), and an exact second moment for the router (<0.01%). The resulting state occupies 1.29 GB, 2.6% of AdamW's, and peak training memory falls from 81.4 GB to 31.3 GB, within the budget of a 40 GB accelerator. In a controlled comparison from identical initializations over 82M tokens, SkewAdam reaches validation perplexity 108.4, ahead of AdamW (126.8), Muon (120.2), and Lion (393.7), and settles router load balance to within 1% of its uniform floor. The allocation is not what earns that perplexity: a tier ablation matches it with twenty times the state, and Adafactor, which shares the factored estimator but drops momentum, plateaus 40 points behind. The tiers buy memory at no cost to accuracy; the accuracy comes from keeping momentum, which a uniform optimizer shares too. Sweeping the baselines' learning rates narrows but does not close the gap: the best tuned AdamW reaches 118.5, tuned Adafactor 139.7. Where optimizer state lives, these results suggest, matters at least as much as how much of it there is.

Read the original paper

More in Optimization

Browse all 36 papers →