NTH

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

AuthorsNikhil Nayak, Julia White, Urchade Zaratiana, Kelton Zhang, Henrijs Princis, Dhruv Atreja, Henry Fawcett, Matthew Thomas, George Hurn-Maloney, Ash Lewis

May 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that popular optimizer updates for language models are subtly biased on small batches, and it introduces a practical correction that can make AdamW, Sophia, and Shampoo train a bit better.

Key results

256K packed FineWeb-Edu sequences
Qwen2.5-0.5B pretraining corpus

Clean pretraining was run from random initialization on Qwen2.5-0.5B using 256K packed FineWeb-Edu training sequences.

0.1489 nats
AdamW held-out loss reduction

The corrected AdamW variant (LOO+Jensen BC) reduced clean held-out pretraining cross-entropy from 4.836 to 4.687.

0.0701 nats
Sophia-G held-out loss reduction

Sophia-G full bias correction improved clean held-out pretraining loss from 6.6647 to 6.5946.

0.1103 nats
Shampoo held-out loss reduction

Shampoo full bias correction improved clean held-out pretraining loss from 5.7916 to 5.6813.

0.0191 nats
Mixed-quality pretraining gain

With 20 percent span-replaced training sequences, corrected AdamW still improved clean held-out loss from 4.8225 to 4.8034.

32K supervised examples
Instruction tuning training set

Instruction tuning used a 32K-example Alpaca-style supervised training subset; effects were mostly neutral with small gains in selected AdamW and Sophia settings.

What the paper found

This paper shows that stochastic preconditioned language-model optimizers such as AdamW, Sophia-G, and Shampoo are biased estimators of the ideal population update because the gradient and preconditioner are usually computed from the same minibatch, and because matrix or scalar inversion is nonlinear. The authors separate these effects into gradient–preconditioner coupling bias and inverse-preconditioner bias, then correct both with a single-batch framework: cross-fitting estimates the numerator gradient and denominator preconditioner from independent microbatch groups, and a delta-method variance correction subtracts the leading inverse bias using microbatch variability. For AdamW and Sophia, the correction is applied elementwise to diagonal preconditioners; for Shampoo, it is applied in the eigenspace of the averaged matrix preconditioner. On Qwen2.5-0.5B pretraining over 256K packed FineWeb-Edu sequences, the corrected optimizers reduce held-out cross-entropy by 0.1489 nats for AdamW using a leave-one-out plus Jensen variant, 0.0701 nats for Sophia-G, and 0.1103 nats for Shampoo. In a mixed-quality pretraining setup with 20 percent span-replaced training sequences, AdamW still improves by 0.0191 nats on clean held-out data despite higher training loss, suggesting reduced overfitting to corrupted spans. Instruction tuning on 32K Alpaca-style examples is mostly neutral, with small gains for retuned AdamW and near-ties for Sophia and Shampoo. The main novelty is not a new optimizer geometry, but a statistical bias-correction layer that makes existing preconditioned updates closer to their population target.

Original abstract

Preconditioned optimizers are central to language model training, but their stochastic update rules are usually treated as direct approximations to population preconditioned descent. We show that this view misses two finite-sample biases. First, the gradient and preconditioner are typically estimated from the same minibatch, introducing gradient--preconditioner coupling bias. Second, even when the preconditioner estimate is unbiased, its inverse or inverse-root is generally biased because inversion is nonlinear. We propose a single-batch bias-correction framework that addresses both effects: cross-fitted preconditioning estimates the numerator and preconditioner from independent microbatch groups, while variance-corrected inversion uses microbatch variability to subtract the leading delta-method bias term. The framework applies to diagonal moment, diagonal curvature, and matrix preconditioning methods, instantiated in AdamW, Sophia, and Shampoo. Bias correction reduces held-out pretraining loss on Qwen2.5-0.5B by $0.15$, $0.07$, and $0.11$ nats, respectively; the effects on mixed-quality pretraining and downstream instruction tuning are consistently neutral-to-positive. Together, these results establish bias correction as a practical mechanism for reducing finite-sample update bias and improving the performance of preconditioned optimizers.

Read the original paper

More in Optimization

Browse all 36 papers →