NTH

Why Muon Outperforms Adam: A Curvature Perspective

AuthorsShuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann, Zhuoran Yang

July 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains why the Muon optimizer can beat Adam in large language model training by showing it makes smaller curvature-induced mistakes, especially in imbalanced and high-curvature settings.

Key results

124M
FineWeb model size

NanoGPT scale used for the main curvature experiments

9M
Zipf-PCFG model size

NanoGPT scale used for the synthetic imbalance experiments

1.76
NDS ratio

Average Adam-to-Muon NDS ratio at matched validation loss

1.8x
NDS gap increase

Widening of the Adam–Muon NDS gap from s=0 to s=1

14%
Muon within-layer share

Early-training within-layer fraction of NDS for Muon

44%
Muon within-layer share

Late-training within-layer fraction of NDS for Muon

What the paper found

This paper, from National University of Singapore, Yale University, and the University of Minnesota, gives a curvature-based explanation for why Muon outperforms Adam in large language model pretraining. Using a 124M-parameter NanoGPT trained on FineWeb-10B, the authors show that at matched validation loss Muon achieves a larger one-step loss decrease because it pays a much smaller second-order curvature penalty while matching Adam’s first-order gain. They define this penalty through normalized directional sharpness (NDS), and find that Muon and Adam have comparable update norms, so the advantage comes from a 1.76× lower Adam-to-Muon NDS ratio, not smaller step size. On synthetic Zipf-PCFG data with a 9M-parameter NanoGPT, increasing imbalance from s=0 to s=1 widens the NDS gap from 0.63 to 1.13, a 1.8× increase, showing that heavy-tailed training data amplifies Muon’s geometric advantage. A layerwise decomposition further shows Muon’s NDS shifts toward within-layer Hessian blocks over training, with the within-layer share rising from 14% to 44%, while Adam stays near 30%. The theory section proves that on structured quadratic models with low Kronecker-rank Hessians and aligned gradients, spectral normalization spreads update energy across curvature modes, yielding smaller average NDS and, under strong curvature heterogeneity, a lower loss than gradient descent after the same number of steps.

Original abstract

Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a first step toward demystifying Muon's superiority over Adam from a curvature perspective. First, we apply a second-order Taylor approximation to the training landscape and show that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. The two optimizers have comparable first-order gains, but Muon consistently incurs a smaller second-order curvature penalty. Second, we decompose this curvature penalty into the squared update norm and Normalized Directional Sharpness (NDS). We find that Muon and Adam have comparable update norms, so Muon's smaller curvature penalty is driven by lower NDS, not update scale. Third, we study how training data and model structure shape Muon's NDS advantage. Using Zipf-Probabilistic Context-Free Grammar (PCFG) data with controlled imbalance, we show that data imbalance amplifies Muon's NDS advantage over Adam. A within-/cross-layer decomposition further shows that, in the middle and late stages of training, Muon's lower NDS is mainly sustained by smaller within-layer curvature. Beyond empirical evidence, we analyze stylized quadratic problems with heterogeneous curvature and gradient alignment toward high-curvature modes. We prove that Muon attains a smaller average NDS than GD by balancing update energy across curvature groups; when curvature heterogeneity is sufficiently strong, this also yields lower local quadratic loss after the same number of steps.

Read the original paper

More in Optimization

Browse all 36 papers →