Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
AuthorsWenzhi Zhong, Edward Milsom, Michael Murray
Resources
This work combines spectral-norm sharpness control with the Muon optimizer to improve ImageNet training performance on both ViTs and ResNets.
Key results
Best ImageNet-1K Top-1 validation accuracy among evaluated methods.
Improvement from the non-SAM Muon baseline at 74.95 to SpecSAM-Muon.
Best ImageNet-1K Top-1 validation accuracy, versus 77.05 for non-SAM Muon.
Validation-accuracy gain produced by the spectral inner step with Muon.
Top-1 accuracy on the out-of-distribution artistic-rendition benchmark.
What the paper found
This University of Bath study examines whether Sharpness-Aware Minimization should measure parameter perturbations with matrix-aware spectral geometry rather than a flattened Euclidean norm. The authors introduce SpecSAM, which applies independent layerwise spectral-norm perturbations to weight matrices, approximating orthogonalization with Newton–Schulz iterations, and pairs this inner step with either AdamW, SGDW, or Muon—the momentum-plus-orthogonalization optimizer increasingly used in frontier language-model systems such as DeepSeek. On ImageNet-1K, SpecSAM-Muon achieves the strongest validation results for both ViT-Small/16 and ResNet-50: 80.23 Top-1 accuracy for the transformer, a 5.28% improvement over the non-SAM Muon baseline at 74.95, and 78.55 for ResNet-50, a 1.50-point gain over Muon at 77.05. The same method reaches 42.18 accuracy on the out-of-distribution ImageNet-R benchmark, although the three-seed differences there are not statistically resolvable. The central finding is an interaction between inner and outer geometry: Euclidean SAM adds roughly similar gains to different optimizers, while spectral perturbations selectively amplify Muon. A ResNet-50 ablation shows that layerwise budgeting alone is insufficient; replacing spectral norms with Frobenius norms removes the Muon advantage. CIFAR-100 diagnostics further indicate that Muon preserves higher effective weight and feature ranks across perturbation radii, potentially retaining representational capacity under stronger regularization. The evidence remains limited to image classification, three-seed main experiments, and methods that require two backward passes plus Newton–Schulz overhead.
Original abstract
Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study matrix-aware geometry in both stages of SAM: we introduce a layerwise spectral inner perturbation for matrix-valued hidden-layer parameters and combine it with either AdamW/SGDW or Muon in the outer update. Across ImageNet-1K experiments on ViT-Small/16 and ResNet-50, we find that the combination of a spectral inner step with a Muon outer step performs consistently strongly, achieving the best validation accuracy on both models among the evaluated methods.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.