On the Principles Behind Neural Network Optimizers
AuthorsYushun Zhang
Resources
This work explains why Adam works so well for modern neural networks and proposes a slimmer version that retains its performance.
Key results
Reduction in optimizer-state memory while preserving Adam-like performance
Largest model size in the Llama 2 experiments on C4
Potential times-faster rate than gradient descent under heterogeneous block structure
Approximate memory required by Adam’s first- and second-moment states
What the paper found
This research develops a theoretical foundation for neural-network optimizers, focusing on Adam, the method used widely to train models such as ChatGPT, Llama, and DeepSeek. It resolves part of Adam’s divergence debate by showing a problem- and batch-size-dependent phase transition: with sufficiently large beta2, unmodified Adam converges toward critical points, while small-beta2 settings can make iterates, gradients, and losses diverge. The explanation for Adam’s advantage over SGD on Transformers is geometric: training drives their Hessians toward near-block-diagonal structure with strongly heterogeneous block spectra, allowing Adam’s diagonal preconditioner to assign effectively different learning rates across blocks. Random-matrix analysis attributes this structure primarily to consecutive multiplications of large weight matrices, rather than specifically to cross-entropy loss. These findings also clarify why Adam is less advantageous on CNNs and dense quadratic problems, and motivate head-wise Muon variants for Transformer attention. The practical contribution is Adam-mini, which groups parameters by Hessian-aligned rows or attention heads and replaces most coordinate-level second-moment statistics with block means; it reduces optimizer-state memory by 50% while matching or sometimes exceeding AdamW. In Llama 2 pretraining on C4, experiments covered models from 39M to 1B parameters. The thesis also estimates that Adam could be 5 times faster than gradient descent on heterogeneous block-diagonal quadratics, while noting that Adam’s states alone require about 104 GB for a 13B model.
Original abstract
Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a principled grounding for Adam and motivates new designs. First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batch-size-dependent hyperparameters, Adam converges, whereas under small-$β_2$ regimes it can diverge. Second, we investigate why Adam substantially outperforms SGD on Transformers through Hessian structure. We find that the Hessian evolves toward a near-block-diagonal form along training, accompanied by strong block heterogeneity. We prove that this structure makes Adam's diagonal preconditioner effective. We further show that this special Hessian structure originates from consecutive multiplications of large matrix variables, and we provide a rigorous analysis based on random matrix theory. Finally, these insights motivate Adam-mini, a new optimizer that reduces Adam's memory footprint by 50\% while preserving its performance. Our results also have broader implications beyond Adam: they reveal new local structures in matrix-based nonconvex problems, and also help understand and improve recent NN optimizers, such as Muon.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.