NTH

On the Principles Behind Neural Network Optimizers

AuthorsYushun Zhang

August 24, 2026 2 min read
Watch on YouTube
The one-line take

This work explains why Adam works so well for modern neural networks and proposes a slimmer version that retains its performance.

Key results

50%
Adam-mini memory reduction

Reduction in optimizer-state memory while preserving Adam-like performance

1B
Llama 2 pretraining scale

Largest model size in the Llama 2 experiments on C4

5
Adam estimated speedup

Potential times-faster rate than gradient descent under heterogeneous block structure

104 GB
Adam state memory for 13B

Approximate memory required by Adam’s first- and second-moment states

What the paper found

This research develops a theoretical foundation for neural-network optimizers, focusing on Adam, the method used widely to train models such as ChatGPT, Llama, and DeepSeek. It resolves part of Adam’s divergence debate by showing a problem- and batch-size-dependent phase transition: with sufficiently large beta2, unmodified Adam converges toward critical points, while small-beta2 settings can make iterates, gradients, and losses diverge. The explanation for Adam’s advantage over SGD on Transformers is geometric: training drives their Hessians toward near-block-diagonal structure with strongly heterogeneous block spectra, allowing Adam’s diagonal preconditioner to assign effectively different learning rates across blocks. Random-matrix analysis attributes this structure primarily to consecutive multiplications of large weight matrices, rather than specifically to cross-entropy loss. These findings also clarify why Adam is less advantageous on CNNs and dense quadratic problems, and motivate head-wise Muon variants for Transformer attention. The practical contribution is Adam-mini, which groups parameters by Hessian-aligned rows or attention heads and replaces most coordinate-level second-moment statistics with block means; it reduces optimizer-state memory by 50% while matching or sometimes exceeding AdamW. In Llama 2 pretraining on C4, experiments covered models from 39M to 1B parameters. The thesis also estimates that Adam could be 5 times faster than gradient descent on heterogeneous block-diagonal quadratics, while noting that Adam’s states alone require about 104 GB for a 13B model.

Original abstract

Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a principled grounding for Adam and motivates new designs. First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batch-size-dependent hyperparameters, Adam converges, whereas under small-$β_2$ regimes it can diverge. Second, we investigate why Adam substantially outperforms SGD on Transformers through Hessian structure. We find that the Hessian evolves toward a near-block-diagonal form along training, accompanied by strong block heterogeneity. We prove that this structure makes Adam's diagonal preconditioner effective. We further show that this special Hessian structure originates from consecutive multiplications of large matrix variables, and we provide a rigorous analysis based on random matrix theory. Finally, these insights motivate Adam-mini, a new optimizer that reduces Adam's memory footprint by 50\% while preserving its performance. Our results also have broader implications beyond Adam: they reveal new local structures in matrix-based nonconvex problems, and also help understand and improve recent NN optimizers, such as Muon.

Read the original paper

More in Optimization

Browse all 36 papers →