NTH

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

AuthorsXiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland

September 1, 2026 3 min read
Watch on YouTube
The one-line take

This work argues that Muon wins by allocating step sizes across spectral directions and uses that insight to make training even more token-efficient.

Key results

13.3%–24.0%
SAMuon token-efficiency gain

Fewer training tokens than Muon to reach the same validation loss.

13.3%–22.1%
SAMuon-lite token-efficiency gain

Token reduction relative to Muon across model-scale and batch-size settings.

124M–1B
Benchmark model range

Parameter range of the modded-nanogpt models used for training experiments.

100B
FineWeb sample size

Token sample used for pretraining.

0.5%
SAMuon-lite overhead

Measured wall-clock overhead over Muon on the 1B benchmark.

7.4%
SAMuon overhead

Measured wall-clock overhead of the current randomized-SVD implementation.

What the paper found

This paper explains Muon’s advantage over Adam and SGD through out-of-sample spectral probing of Transformer momentum buffers. In GPT-2-style decoder-only models, including architectures related to OpenAI’s GPT-2 rather than ChatGPT itself, the loss-optimal step profile is strongly anisotropic: a dominant volatile head sits at the Edge-of-Stability and limits the global learning rate, while a tolerant bulk supports substantially larger updates. SGD over-allocates to the head through singular-value-proportional scaling, Adam partially dampens that mismatch, and Muon’s whitening reallocates update strength toward the bulk, explaining the observed ranking Muon over Adam over SGD. The paper then introduces Spectral-Aware Muon, or SAMuon, which preserves the head at Muon’s scale while amplifying the bulk using a static spectral prior. Full SAMuon follows the approximately log-rank-linear profile with randomized low-rank SVD; SAMuon-lite uses rank-one power iteration and a two-level head-versus-bulk allocation. On FineWeb’s 100B-token sample, “modded-nanogpt” models from 124M to 1B parameters show that SAMuon reaches Muon’s validation loss with 13.3%–24.0% fewer tokens, while SAMuon-lite achieves 13.3%–22.1%. SAMuon-lite adds only 0.5% wall-clock overhead, compared with 7.4% for the current full-SVD implementation. The analysis also links unstable attention-tail directions to QK clipping practices reported for Kimi K2, suggesting spectral allocation can guide optimizer and architecture design.

Original abstract

Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. At checkpoints along real training trajectories, we decompose each momentum buffer into its singular directions and estimate the loss-optimal step size along each direction on held-out data. The resulting spectral profile is anisotropic yet stable across batches and training stages, and consistent across the optimisers and model scales: a volatile head operating at the Edge-of-Stability supports a much smaller step size than the tolerant bulk, which permits substantially larger steps. This profile provides a unified spectral allocation account of why Muon outperforms Adam, which outperforms SGD. It also exposes a limitation of Muon's uniform scaling: it still underutilises the bulk. Guided by this finding, we introduce Spectral-Aware Muon (SAMuon), which holds the head at the Muon scale and amplifies the bulk using a static spectral prior. We provide two variants: the complete SAMuon follows the measured profile using a low-rank randomised SVD and the simplified SAMuon-lite uses a two-level approximation via rank-one power iteration. Neither method adds persistent optimiser state or notable extra FLOPs beyond Muon at scale, and the idealised exact-whitening versions of both retain Muon's asymptotic convergence rate under standard assumptions. Across "modded-nanogpt" models from 124M to 1B parameters, both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations. SAMuon requires 13.3% to 24.0% fewer training tokens to reach the same validation loss as Muon, while SAMuon-lite retains most of this gain with near-zero wall-clock overhead.

Read the original paper

More in Optimization

Browse all 36 papers →