Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
AuthorsXiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland
Resources
This work argues that Muon wins by allocating step sizes across spectral directions and uses that insight to make training even more token-efficient.
Key results
Fewer training tokens than Muon to reach the same validation loss.
Token reduction relative to Muon across model-scale and batch-size settings.
Parameter range of the modded-nanogpt models used for training experiments.
Token sample used for pretraining.
Measured wall-clock overhead over Muon on the 1B benchmark.
Measured wall-clock overhead of the current randomized-SVD implementation.
What the paper found
This paper explains Muon’s advantage over Adam and SGD through out-of-sample spectral probing of Transformer momentum buffers. In GPT-2-style decoder-only models, including architectures related to OpenAI’s GPT-2 rather than ChatGPT itself, the loss-optimal step profile is strongly anisotropic: a dominant volatile head sits at the Edge-of-Stability and limits the global learning rate, while a tolerant bulk supports substantially larger updates. SGD over-allocates to the head through singular-value-proportional scaling, Adam partially dampens that mismatch, and Muon’s whitening reallocates update strength toward the bulk, explaining the observed ranking Muon over Adam over SGD. The paper then introduces Spectral-Aware Muon, or SAMuon, which preserves the head at Muon’s scale while amplifying the bulk using a static spectral prior. Full SAMuon follows the approximately log-rank-linear profile with randomized low-rank SVD; SAMuon-lite uses rank-one power iteration and a two-level head-versus-bulk allocation. On FineWeb’s 100B-token sample, “modded-nanogpt” models from 124M to 1B parameters show that SAMuon reaches Muon’s validation loss with 13.3%–24.0% fewer tokens, while SAMuon-lite achieves 13.3%–22.1%. SAMuon-lite adds only 0.5% wall-clock overhead, compared with 7.4% for the current full-SVD implementation. The analysis also links unstable attention-tail directions to QK clipping practices reported for Kimi K2, suggesting spectral allocation can guide optimizer and architecture design.
Original abstract
Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. At checkpoints along real training trajectories, we decompose each momentum buffer into its singular directions and estimate the loss-optimal step size along each direction on held-out data. The resulting spectral profile is anisotropic yet stable across batches and training stages, and consistent across the optimisers and model scales: a volatile head operating at the Edge-of-Stability supports a much smaller step size than the tolerant bulk, which permits substantially larger steps. This profile provides a unified spectral allocation account of why Muon outperforms Adam, which outperforms SGD. It also exposes a limitation of Muon's uniform scaling: it still underutilises the bulk. Guided by this finding, we introduce Spectral-Aware Muon (SAMuon), which holds the head at the Muon scale and amplifies the bulk using a static spectral prior. We provide two variants: the complete SAMuon follows the measured profile using a low-rank randomised SVD and the simplified SAMuon-lite uses a two-level approximation via rank-one power iteration. Neither method adds persistent optimiser state or notable extra FLOPs beyond Muon at scale, and the idealised exact-whitening versions of both retain Muon's asymptotic convergence rate under standard assumptions. Across "modded-nanogpt" models from 124M to 1B parameters, both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations. SAMuon requires 13.3% to 24.0% fewer training tokens to reach the same validation loss as Muon, while SAMuon-lite retains most of this gain with near-zero wall-clock overhead.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.