The Loss Does Not See the Basis, but Adam Does
AuthorsDevender Singh
Resources
The paper argues that Adam’s coordinate-wise behavior breaks a hidden symmetry that lets gradient descent favor low-rank solutions, showing that optimizer geometry—not just the loss—determines which answer a model learns.
Key results
Ground-truth recovery error for standard Adam on the 40×40 rank-3 matrix-sensing task.
Ground-truth recovery error for gradient descent under the same interpolation protocol.
Near-exact recovery achieved by Muon on the exactly low-rank sensing target.
Relative Frobenius gap between Adam-trained, gauge-equivalent per-head WQᵀWK invariants.
Gradient descent’s held-out error reduction versus Adam on Indian Pines at the lowest sampling density.
Approximate target tail-energy boundary where Muon loses its advantage over gradient descent.
What the paper found
The paper argues that a factored model W=UVᵀ has a gauge symmetry: rotating both latent factors by the same orthogonal basis leaves the loss and predictions unchanged, but optimizers do not all respect that symmetry. Gradient descent, momentum, shared-scalar Adam, Muon, and Shampoo are gauge-equivariant, while standard Adam, RMSProp, Lion, signSGD, and Adafactor use coordinate-wise statistics that make the chosen basis affect which interpolating solution is learned. On a 40×40 rank-3 matrix-sensing problem, all nine methods reached essentially zero training residual, yet Adam’s recovery error was 0.5734 versus 0.1312 for gradient descent, showing that the difference is implicit solution selection rather than fit quality; Muon achieved near-exact recovery at 0.0000068. A continuous Adam-p dial isolates coordinate anisotropy: moving from p=1, standard Adam, to p=0, shared-scalar Adam, monotonically lowers effective rank and improves recovery. The same issue appears inside transformer attention heads, where WQᵀWK is gauge-invariant: on modular addition modulo 47, Adam’s basis-equivalent twins diverged at the first step and their final per-head invariants differed by 56%, a discrepancy no rotation can remove. Spectral scheduling is a second axis: Muon excels on exactly low-rank targets but loses its advantage once spectral-tail energy reaches about 4%. On Indian Pines and Pavia University hyperspectral completion, gradient descent reduced held-out error by 44% at matched training loss in the most underdetermined setting. The findings also matter for LoRA-style factorizations and model merging: preserving internal symmetry can determine whether two functionally equivalent initializations remain compatible.
Original abstract
Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mechanism is available to an optimizer only if that optimizer is gauge-equivariant, a condition necessary for the transfer but not sufficient for low-rank recovery. Gradient descent, momentum, "shared-scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate-wise methods do not. A structure theorem characterizes the memoryless equivariant rules as exactly the Gram-determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common-scalar flows. We then sort nine update rules on underdetermined matrix sensing by recovery error against the planted ground truth. A one-parameter family from coordinate-wise to shared-scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. A "spectral schedule" reconciles two opposing reports about Muon: equal-rate updates recover exactly low-rank targets but lose their edge as the spectral tail grows. In transformers, Adam separates two gauge-equivalent initializations at the first step, where the equivariant optimizers stay at float precision, and ends with the per-head invariants $W_Q^\top W_K$ 56% apart in relative Frobenius distance, a gap no per-head rotation can close. On two hyperspectral datasets at matched training loss, gradient descent cuts held-out error by 43-44% at the lowest sampling density, and at lower effective rank. Basis choice is therefore not a tuning detail but a decision about which interpolant the optimizer selects.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.