Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
AuthorsZihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, Zhiqiang Zhang, Lei Wu
Resources
The paper argues that language-model training dynamics are governed less by raw learning rate than by its ratio to parameter norm, revealing a common effective-learning-rate scale.
Key results
Median loss discrepancy across 26 ELR-matched comparisons
Error after removing QK-Norm and fixing RMSNorm gains in Llama-124M on FineWeb
Loss alignment between AdamW runs with and without weight decay
Loss alignment between Hyperball and weight-decay Muon runs
Prediction error on unseen Hyperball runs without refitting
Fold reduction in RMSE versus learning-rate-based FSL on Hyperball
What the paper found
This paper identifies the effective learning rate, or ELR, defined as the learning rate divided by the parameter Frobenius norm, as the main coordinate governing language-model pretraining loss dynamics. When ELR schedules are matched, runs with substantially different learning-rate and norm trajectories produce nearly identical losses: across 26 comparisons spanning Llama, Qwen3, Kimi Delta Attention, FineWeb, C4, OpenWebText, AdamW, Muon, and Signum, the median collapse error is 2.5e-3 and every comparison stays below 5e-3. The effect is conditional rather than a consequence of exact scale invariance: in Llama-124M on FineWeb, removing QK-Norm raises error from 2.3e-3 to 5.2e-3, while also fixing RMSNorm gains increases it to 1.84e-2. Practical norm-control methods behave through the ELR schedules they induce; matching ELR aligns weight-decay and Hyperball training with errors of 4.8e-3 and 1.2e-3. Replacing learning rate with ELR in a functional scaling law enables transfer to unseen Hyperball runs, reducing RMSE to 0.0212 instead of 0.2508, an 11.83-fold improvement. The framework also explains delayed acceleration: norm control can preserve a larger ELR early, improving effective training, while late ELR decay allows accumulated noise to diminish and reveals the gain. The practical recommendation is to design the ELR schedule first, then select a learning-rate and norm-control mechanism that realizes it.
Original abstract
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce. Replacing LR with ELR enables a fitted functional scaling law (FSL) to transfer across norm-control methods. The resulting ELR-based FSL also explains delayed acceleration, a recurring effect of norm control. Together, these results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.