Approximate Muon with low-rank adapters
AuthorsBen Anson, Conor Houghton, Edward Milsom
Resources
sMuon adapts the Muon optimizer to low-rank fine-tuning, offering a simpler and sometimes more effective alternative for LoRA-based training.
Key results
sMuon achieved the top accuracy on 6 of 11 SFT tasks with Moonlight-16B-A3B.
sMuon scored 86.2% on PIQA with Moonlight-16B-A3B.
sMuon tied Riemannion for the best final validation loss.
sMuon was approximately 30% faster than Riemannion.
The sMuon optimizer step took 12 milliseconds at LoRA rank 64.
The ReLoRA experiment trained a 162M-parameter transformer on FineWeb.
What the paper found
This paper introduces sMuon, an optimizer that brings Muon’s spectral-norm steepest-descent geometry to LoRA adapters, where directly orthogonalizing the full update is incompatible with the low-rank form ΔW = BA. sMuon linearizes the Muon objective, projects the gradient into the adapter’s realizable subspace, and solves a least-squares update for the two factors, with momentum transport and split weight decay preserving parameterization invariance. Its efficient implementation uses matrix multiplications, inverse roots, and a matrix-sign operation on a 2r-by-2r core, avoiding the QR and SVD decompositions required by Riemannion. Across 4 models—Qwen2.5-3B, Llama-3.2-3B, DeepSeek-V2-Lite, and the Muon-pretrained Moonlight-16B-A3B—over 11 SFT benchmarks, sMuon achieved the best result on 6 tasks with Moonlight, including 86.2% on PIQA. In ReLoRA pretraining of a 162M-parameter transformer on FineWeb, sMuon tied Riemannion for the best final validation loss at 3.66 while running approximately 30% faster. At rank 64, its optimizer step took 12 milliseconds, compared with 700 milliseconds for Riemannion, demonstrating that geometry-aware low-rank Muon updates can retain much of the optimization benefit without expensive manifold retractions.
Original abstract
The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank parameterization. In this paper, we address this issue by approximating the solution to a relaxed Muon objective in the low-rank setting via linearization and then least-squares. We provide an efficient implementation that uses matmul operations only, as opposed to more complex linear algebra decomposition routines. Our method, sMuon (small Muon), performs favourably across SFT and a ReLoRA pretraining experiment. While results are model- and eval-dependent, we find overall that using Muon for low-rank fine-tuning provides moderate performance improvements.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.