NTH

Approximate Muon with low-rank adapters

AuthorsBen Anson, Conor Houghton, Edward Milsom

August 24, 2026 2 min read
Watch on YouTube
The one-line take

sMuon adapts the Muon optimizer to low-rank fine-tuning, offering a simpler and sometimes more effective alternative for LoRA-based training.

Key results

6
Moonlight task wins

sMuon achieved the top accuracy on 6 of 11 SFT tasks with Moonlight-16B-A3B.

86.2%
Moonlight PIQA accuracy

sMuon scored 86.2% on PIQA with Moonlight-16B-A3B.

3.66
ReLoRA final validation loss

sMuon tied Riemannion for the best final validation loss.

30%
ReLoRA speed advantage

sMuon was approximately 30% faster than Riemannion.

12
sMuon step time at rank 64

The sMuon optimizer step took 12 milliseconds at LoRA rank 64.

162M
ReLoRA model size

The ReLoRA experiment trained a 162M-parameter transformer on FineWeb.

What the paper found

This paper introduces sMuon, an optimizer that brings Muon’s spectral-norm steepest-descent geometry to LoRA adapters, where directly orthogonalizing the full update is incompatible with the low-rank form ΔW = BA. sMuon linearizes the Muon objective, projects the gradient into the adapter’s realizable subspace, and solves a least-squares update for the two factors, with momentum transport and split weight decay preserving parameterization invariance. Its efficient implementation uses matrix multiplications, inverse roots, and a matrix-sign operation on a 2r-by-2r core, avoiding the QR and SVD decompositions required by Riemannion. Across 4 models—Qwen2.5-3B, Llama-3.2-3B, DeepSeek-V2-Lite, and the Muon-pretrained Moonlight-16B-A3B—over 11 SFT benchmarks, sMuon achieved the best result on 6 tasks with Moonlight, including 86.2% on PIQA. In ReLoRA pretraining of a 162M-parameter transformer on FineWeb, sMuon tied Riemannion for the best final validation loss at 3.66 while running approximately 30% faster. At rank 64, its optimizer step took 12 milliseconds, compared with 700 milliseconds for Riemannion, demonstrating that geometry-aware low-rank Muon updates can retain much of the optimization benefit without expensive manifold retractions.

Original abstract

The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank parameterization. In this paper, we address this issue by approximating the solution to a relaxed Muon objective in the low-rank setting via linearization and then least-squares. We provide an efficient implementation that uses matmul operations only, as opposed to more complex linear algebra decomposition routines. Our method, sMuon (small Muon), performs favourably across SFT and a ReLoRA pretraining experiment. While results are model- and eval-dependent, we find overall that using Muon for low-rank fine-tuning provides moderate performance improvements.

Read the original paper

More in Optimization

Browse all 36 papers →