NTH

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

AuthorsAratrika Mustafi, Soumya Mukherjee, Bharath K. Sriperumbudur

May 29, 2026 3 min read
Watch on YouTube
The one-line take

This paper gives Muon optimizer a physics-style mathematical makeover, showing it as a damped Hamiltonian flow and deriving convergence and mean-field guarantees.

Key results

2.0e-32
matrix mean matching (1,10) best regularized Muon

Final objective value in the deterministic matrix mean-matching experiment for the (M,N)=(1,10) setting.

1.3e-7
teacher-student (3,12) best regularized Muon

Final objective value in the nonlinear product-space teacher-student experiment for the (M,N)=(3,12) setting.

6.0e-7
teacher-student (10,10) best regularized Muon

Final objective value in the nonlinear product-space teacher-student experiment for the (M,N)=(10,10) setting.

1.1e-3
matrix mean matching residual floor, hard Muon

Residual objective floor reported for hard Muon in both matrix mean-matching settings.

What the paper found

This paper reframes Muon as a spectral trust-region mirror method and then lifts it to a phase-space probability flow on matrix-valued parameters. The central technical move is a smooth regularization of the hard polar-factor map, Orth(P), via Orthε(P)=U diag(σi/√(σi2+ε2))V⊤, which is exactly the gradient of a Fenchel-dual smoothing of the nuclear norm; this makes the Muon step a mirror/prox update with momentum as the dual coordinate rather than an ad hoc normalization. From that structure, the authors derive an inertial continuous-time limit for finite particle systems, then a McKean-Vlasov continuity equation over laws µt on (W,P) pairs. The resulting dynamics are a damped Hamiltonian probability flow with Hamiltonian Hε,γ(µ)=∫Ψε(P)dµ+γR(∫F(W)dµ), and they prove an exact dissipation identity dHε,γ/dt=−γ∫⟨P,Orthε(P)⟩F dµ≤0, even though the target objective J(ρt) itself need not be monotone. Under gradient-dominance, bounded-momentum, and curvature/alignment assumptions, they obtain exponential objective-gap decay; on the scalar specialization, the rate is exp(−cα,rt) with explicit constants. They also prove propagation of chaos with O(1/N) mean-square particle error and show that the hard Muon limit ε→0 yields a subsequential differential inclusion with nuclear-norm subgradients. The framework extends to product matrix spaces and transformer mixture-of-experts models by treating expert weights and router parameters jointly; in synthetic matrix mean-matching and teacher-student experiments, regularized Muon removes the ≈10−3 residual floor seen with hard polar and Newton-Schulz updates and reaches losses as low as 2.0×10−32 in one setting and 6.0×10−7 in another.

Original abstract

We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(ρ)=R\left(\int F d ρ\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.

Read the original paper

More in Optimization

Browse all 36 papers →