Why Muon Outperforms Adam: A Curvature Perspective
AuthorsShuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann, Zhuoran Yang
Resources
This paper explains why the Muon optimizer can beat Adam in large language model training by showing it makes smaller curvature-induced mistakes, especially in imbalanced and high-curvature settings.
Key results
NanoGPT scale used for the main curvature experiments
NanoGPT scale used for the synthetic imbalance experiments
Average Adam-to-Muon NDS ratio at matched validation loss
Widening of the Adam–Muon NDS gap from s=0 to s=1
Early-training within-layer fraction of NDS for Muon
Late-training within-layer fraction of NDS for Muon
What the paper found
This paper, from National University of Singapore, Yale University, and the University of Minnesota, gives a curvature-based explanation for why Muon outperforms Adam in large language model pretraining. Using a 124M-parameter NanoGPT trained on FineWeb-10B, the authors show that at matched validation loss Muon achieves a larger one-step loss decrease because it pays a much smaller second-order curvature penalty while matching Adam’s first-order gain. They define this penalty through normalized directional sharpness (NDS), and find that Muon and Adam have comparable update norms, so the advantage comes from a 1.76× lower Adam-to-Muon NDS ratio, not smaller step size. On synthetic Zipf-PCFG data with a 9M-parameter NanoGPT, increasing imbalance from s=0 to s=1 widens the NDS gap from 0.63 to 1.13, a 1.8× increase, showing that heavy-tailed training data amplifies Muon’s geometric advantage. A layerwise decomposition further shows Muon’s NDS shifts toward within-layer Hessian blocks over training, with the within-layer share rising from 14% to 44%, while Adam stays near 30%. The theory section proves that on structured quadratic models with low Kronecker-rank Hessians and aligned gradients, spectral normalization spreads update energy across curvature modes, yielding smaller average NDS and, under strong curvature heterogeneity, a lower loss than gradient descent after the same number of steps.
Original abstract
Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a first step toward demystifying Muon's superiority over Adam from a curvature perspective. First, we apply a second-order Taylor approximation to the training landscape and show that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. The two optimizers have comparable first-order gains, but Muon consistently incurs a smaller second-order curvature penalty. Second, we decompose this curvature penalty into the squared update norm and Normalized Directional Sharpness (NDS). We find that Muon and Adam have comparable update norms, so Muon's smaller curvature penalty is driven by lower NDS, not update scale. Third, we study how training data and model structure shape Muon's NDS advantage. Using Zipf-Probabilistic Context-Free Grammar (PCFG) data with controlled imbalance, we show that data imbalance amplifies Muon's NDS advantage over Adam. A within-/cross-layer decomposition further shows that, in the middle and late stages of training, Muon's lower NDS is mainly sustained by smaller within-layer curvature. Beyond empirical evidence, we analyze stylized quadratic problems with heterogeneous curvature and gradient alignment toward high-curvature modes. We prove that Muon attains a smaller average NDS than GD by balancing update energy across curvature groups; when curvature heterogeneity is sufficiently strong, this also yields lower local quadratic loss after the same number of steps.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.