NTH

Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

AuthorsWenzhi Zhong, Edward Milsom, Michael Murray

August 7, 2026 2 min read
Watch on YouTube
The one-line take

This work combines spectral-norm sharpness control with the Muon optimizer to improve ImageNet training performance on both ViTs and ResNets.

Key results

80.23
SpecSAM-Muon ViT-Small/16 validation accuracy

Best ImageNet-1K Top-1 validation accuracy among evaluated methods.

5.28%
ViT-Small/16 gain over Muon

Improvement from the non-SAM Muon baseline at 74.95 to SpecSAM-Muon.

78.55
SpecSAM-Muon ResNet-50 validation accuracy

Best ImageNet-1K Top-1 validation accuracy, versus 77.05 for non-SAM Muon.

1.50
ResNet-50 gain over Muon

Validation-accuracy gain produced by the spectral inner step with Muon.

42.18
SpecSAM-Muon ImageNet-R accuracy

Top-1 accuracy on the out-of-distribution artistic-rendition benchmark.

What the paper found

This University of Bath study examines whether Sharpness-Aware Minimization should measure parameter perturbations with matrix-aware spectral geometry rather than a flattened Euclidean norm. The authors introduce SpecSAM, which applies independent layerwise spectral-norm perturbations to weight matrices, approximating orthogonalization with Newton–Schulz iterations, and pairs this inner step with either AdamW, SGDW, or Muon—the momentum-plus-orthogonalization optimizer increasingly used in frontier language-model systems such as DeepSeek. On ImageNet-1K, SpecSAM-Muon achieves the strongest validation results for both ViT-Small/16 and ResNet-50: 80.23 Top-1 accuracy for the transformer, a 5.28% improvement over the non-SAM Muon baseline at 74.95, and 78.55 for ResNet-50, a 1.50-point gain over Muon at 77.05. The same method reaches 42.18 accuracy on the out-of-distribution ImageNet-R benchmark, although the three-seed differences there are not statistically resolvable. The central finding is an interaction between inner and outer geometry: Euclidean SAM adds roughly similar gains to different optimizers, while spectral perturbations selectively amplify Muon. A ResNet-50 ablation shows that layerwise budgeting alone is insufficient; replacing spectral norms with Frobenius norms removes the Muon advantage. CIFAR-100 diagnostics further indicate that Muon preserves higher effective weight and feature ranks across perturbation radii, potentially retaining representational capacity under stronger regularization. The evidence remains limited to image classification, three-seed main experiments, and methods that require two backward passes plus Newton–Schulz overhead.

Original abstract

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study matrix-aware geometry in both stages of SAM: we introduce a layerwise spectral inner perturbation for matrix-valued hidden-layer parameters and combine it with either AdamW/SGDW or Muon in the outer update. Across ImageNet-1K experiments on ViT-Small/16 and ResNet-50, we find that the combination of a spectral inner step with a Muon outer step performs consistently strongly, achieving the best validation accuracy on both models among the evaluated methods.

Read the original paper

More in Optimization

Browse all 36 papers →