Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
AuthorsNikhil Nayak, Julia White, Urchade Zaratiana, Kelton Zhang, Henrijs Princis, Dhruv Atreja, Henry Fawcett, Matthew Thomas, George Hurn-Maloney, Ash Lewis
Resources
This paper shows that popular optimizer updates for language models are subtly biased on small batches, and it introduces a practical correction that can make AdamW, Sophia, and Shampoo train a bit better.
Key results
Clean pretraining was run from random initialization on Qwen2.5-0.5B using 256K packed FineWeb-Edu training sequences.
The corrected AdamW variant (LOO+Jensen BC) reduced clean held-out pretraining cross-entropy from 4.836 to 4.687.
Sophia-G full bias correction improved clean held-out pretraining loss from 6.6647 to 6.5946.
Shampoo full bias correction improved clean held-out pretraining loss from 5.7916 to 5.6813.
With 20 percent span-replaced training sequences, corrected AdamW still improved clean held-out loss from 4.8225 to 4.8034.
Instruction tuning used a 32K-example Alpaca-style supervised training subset; effects were mostly neutral with small gains in selected AdamW and Sophia settings.
What the paper found
This paper shows that stochastic preconditioned language-model optimizers such as AdamW, Sophia-G, and Shampoo are biased estimators of the ideal population update because the gradient and preconditioner are usually computed from the same minibatch, and because matrix or scalar inversion is nonlinear. The authors separate these effects into gradient–preconditioner coupling bias and inverse-preconditioner bias, then correct both with a single-batch framework: cross-fitting estimates the numerator gradient and denominator preconditioner from independent microbatch groups, and a delta-method variance correction subtracts the leading inverse bias using microbatch variability. For AdamW and Sophia, the correction is applied elementwise to diagonal preconditioners; for Shampoo, it is applied in the eigenspace of the averaged matrix preconditioner. On Qwen2.5-0.5B pretraining over 256K packed FineWeb-Edu sequences, the corrected optimizers reduce held-out cross-entropy by 0.1489 nats for AdamW using a leave-one-out plus Jensen variant, 0.0701 nats for Sophia-G, and 0.1103 nats for Shampoo. In a mixed-quality pretraining setup with 20 percent span-replaced training sequences, AdamW still improves by 0.0191 nats on clean held-out data despite higher training loss, suggesting reduced overfitting to corrupted spans. Instruction tuning on 32K Alpaca-style examples is mostly neutral, with small gains for retuned AdamW and near-ties for Sophia and Shampoo. The main novelty is not a new optimizer geometry, but a statistical bias-correction layer that makes existing preconditioned updates closer to their population target.
Original abstract
Preconditioned optimizers are central to language model training, but their stochastic update rules are usually treated as direct approximations to population preconditioned descent. We show that this view misses two finite-sample biases. First, the gradient and preconditioner are typically estimated from the same minibatch, introducing gradient--preconditioner coupling bias. Second, even when the preconditioner estimate is unbiased, its inverse or inverse-root is generally biased because inversion is nonlinear. We propose a single-batch bias-correction framework that addresses both effects: cross-fitted preconditioning estimates the numerator and preconditioner from independent microbatch groups, while variance-corrected inversion uses microbatch variability to subtract the leading delta-method bias term. The framework applies to diagonal moment, diagonal curvature, and matrix preconditioning methods, instantiated in AdamW, Sophia, and Shampoo. Bias correction reduces held-out pretraining loss on Qwen2.5-0.5B by $0.15$, $0.07$, and $0.11$ nats, respectively; the effects on mixed-quality pretraining and downstream instruction tuning are consistently neutral-to-positive. Together, these results establish bias correction as a practical mechanism for reducing finite-sample update bias and improving the performance of preconditioned optimizers.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.