Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer
AuthorsAratrika Mustafi, Soumya Mukherjee, Bharath K. Sriperumbudur
Resources
This paper gives Muon optimizer a physics-style mathematical makeover, showing it as a damped Hamiltonian flow and deriving convergence and mean-field guarantees.
Key results
Final objective value in the deterministic matrix mean-matching experiment for the (M,N)=(1,10) setting.
Final objective value in the nonlinear product-space teacher-student experiment for the (M,N)=(3,12) setting.
Final objective value in the nonlinear product-space teacher-student experiment for the (M,N)=(10,10) setting.
Residual objective floor reported for hard Muon in both matrix mean-matching settings.
What the paper found
This paper reframes Muon as a spectral trust-region mirror method and then lifts it to a phase-space probability flow on matrix-valued parameters. The central technical move is a smooth regularization of the hard polar-factor map, Orth(P), via Orthε(P)=U diag(σi/√(σi2+ε2))V⊤, which is exactly the gradient of a Fenchel-dual smoothing of the nuclear norm; this makes the Muon step a mirror/prox update with momentum as the dual coordinate rather than an ad hoc normalization. From that structure, the authors derive an inertial continuous-time limit for finite particle systems, then a McKean-Vlasov continuity equation over laws µt on (W,P) pairs. The resulting dynamics are a damped Hamiltonian probability flow with Hamiltonian Hε,γ(µ)=∫Ψε(P)dµ+γR(∫F(W)dµ), and they prove an exact dissipation identity dHε,γ/dt=−γ∫⟨P,Orthε(P)⟩F dµ≤0, even though the target objective J(ρt) itself need not be monotone. Under gradient-dominance, bounded-momentum, and curvature/alignment assumptions, they obtain exponential objective-gap decay; on the scalar specialization, the rate is exp(−cα,rt) with explicit constants. They also prove propagation of chaos with O(1/N) mean-square particle error and show that the hard Muon limit ε→0 yields a subsequential differential inclusion with nuclear-norm subgradients. The framework extends to product matrix spaces and transformer mixture-of-experts models by treating expert weights and router parameters jointly; in synthetic matrix mean-matching and teacher-student experiments, regularized Muon removes the ≈10−3 residual floor seen with hard polar and Newton-Schulz updates and reaches losses as low as 2.0×10−32 in one setting and 6.0×10−7 in another.
Original abstract
We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(ρ)=R\left(\int F d ρ\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.