Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics
AuthorsIgor Ignashin, Anna Radovskaya, Andrew Semenov, Egor Lopatin, Stanislav Potapov, Aleksandr Kovalenko, Andrey Veprikov, Aleksandr Shestakov, Andrey Leonidov, Aleksandr Beznosikov
Resources
This paper argues that SGD should be viewed less like Brownian motion and more like motion in a randomly changing landscape, revealing when training trajectories diffuse versus stay trapped.
Key results
The small-scale MNIST validation uses a compact MLP with 386 trainable parameters to enable exact dense Hessian computation.
The theory is validated on an MLP trained on MNIST, where the covariance structure is measured in the mean-Hessian eigenbasis.
A simplified GPT-style model is trained on Shakespeare to test whether the same sharp/flat mode separation appears in language modeling.
The large-scale validation uses a NanoGPT model with 6.6M parameters trained on WikiText-2.
The 6.6M-parameter NanoGPT experiment is trained and evaluated on WikiText-2 for the large-scale variance analysis.
At learning rate 0.01, the standard Langevin approximation underestimates the observed variance plateau in the NanoGPT experiment by about 23%.
What the paper found
This paper argues that stochastic gradient descent should not be modeled as Brownian motion with Gaussian forcing at finite learning rates. Starting from the exact discrete SGD update with sampling with replacement, the authors derive a master equation and a second-order discrete Fokker–Planck expansion, showing that standard Langevin approximations miss terms of order eta squared, including a correction proportional to the squared mean gradient. In a one-dimensional quadratic toy model, this mismatch changes the long-time behavior: the exact discrete dynamics can predict variance growth where the Langevin model predicts a stationary distribution, so the error is qualitative, not just numerical. Near critical points, the theory diagonalizes the covariance dynamics in the mean-Hessian eigenbasis and separates SGD into two regimes: sharp positive-curvature directions confine parameter fluctuations to finite variance, while nearly flat directions remain non-stationary and diffuse, with variance growing roughly linearly in time. Empirically, the paper validates these claims on an MLP with 386 parameters on MNIST, a simplified GPT-style model on Shakespeare, and a 6.6M-parameter NanoGPT trained on WikiText-2. In the NanoGPT experiment, the discrete theory matches the observed variance plateaus across the top 20 Hessian eigendirections, while the standard Langevin prediction underestimates the plateau by about 23 percent at learning rate 0.01. The central novelty is the reinterpretation of SGD as deterministic motion in a fluctuating loss landscape, not a particle driven by external Brownian noise.
Original abstract
Stochastic Gradient Descent (SGD) is commonly modeled as a Langevin process, assuming that minibatch noise acts as Brownian motion. However, this approximation relies on a continuous-time limit and a sqrt(eta) noise scaling that does not match the discrete SGD update at finite learning rate. In this work, we propose an alternative formulation of SGD as deterministic dynamics in a fluctuating loss landscape induced by minibatch sampling. Starting directly from the discrete update, we derive a master equation for the parameter distribution and obtain a discrete Fokker--Planck equation that differs from the standard Langevin form at order eta^2. Using this framework, we analyze SGD dynamics near critical points of the loss. We show that the behavior decomposes along the eigenbasis of the mean Hessian into qualitatively distinct regimes. In particular, nearly-flat directions do not admit a stationary distribution: the variance grows over time, corresponding to effective diffusion along valleys with a coefficient proportional to the learning rate. We provide empirical evidence supporting these predictions on neural network models in computer vision and natural language processing, observing a clear qualitative separation between confined and diffusive modes.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.