One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
AuthorsDi He, Songjun Tu, Keyu Wang, Lu Yin, Shiwei Liu
This paper uses heavy-tail theory to automatically give different Transformer layers different learning rates, helping large language models train faster and perform better.
Key results
On FineWeb pretraining, LLR lowers validation perplexity versus tuned Uniform AdamW for the 60M LLaMA model.
LLR improves validation perplexity over Uniform AdamW on the 135M LLaMA model during FineWeb pretraining.
LLR reduces validation perplexity relative to Uniform AdamW for the 350M LLaMA model.
LLR achieves lower validation perplexity than Uniform AdamW on the 1B LLaMA model.
After pretraining, LLR improves average zero-shot commonsense accuracy across PIQA, SIQA, HellaSwag, WinoGrande, ARC-c, ARC-e, and OBQA.
What the paper found
“One LR Doesn’t Fit All” argues that a single global learning rate is a poor match for Transformer heterogeneity in LLM pretraining, and introduces Layerwise Learning Rate, or LLR, which uses Heavy-Tailed Self-Regularization theory to measure each layer’s empirical spectral density, fit a power law, and compute a PL_Alpha_Hill exponent as a proxy for how “heavy-tailed” and well-trained that layer already is. The rule is simple but novel: layers with lower PL_Alpha_Hill, such as attention projections like Att.q and Att.k, receive smaller learning rates, while layers with higher PL_Alpha_Hill, especially FFN and embedding blocks, receive larger ones; LLaMA embeddings are pinned to the upper LR bound, and LR changes are applied with a soft-switch schedule during only the first 20% of training to avoid spikes and reduce overhead. On FineWeb pretraining, LLR improves validation perplexity over tuned Uniform AdamW from 21.94 to 20.30 on LLaMA-60M, 17.86 to 17.03 on 135M, 12.96 to 12.71 on 350M, and 9.77 to 9.59 on 1B, while also outperforming LARS, LAMB, and the sharpness-based baseline. On LLaMA-1B, zero-shot commonsense accuracy across PIQA, SIQA, HellaSwag, WinoGrande, ARC-c, ARC-e, and OBQA rises from 47.09% to 49.02%. The method transfers to GPT-nano, RoBERTa-base fine-tuning, and Muon optimization, and the authors report up to 1.5× faster convergence with substantially lower tuning overhead than prior layerwise schemes.
Original abstract
Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural heterogeneity of Transformers, potentially limiting their effectiveness as the backbone of Large Language Models (LLMs). In this paper, we introduce Layerwise Learning Rate (LLR), an adaptive scheme that assigns distinct learning rates to individual Transformer layers. Our method is grounded in Heavy-Tailed Self-Regularization (HT-SR) theory, which characterizes the empirical spectral density (ESD) of weight correlation matrices to quantify heavy-tailedness. Layers with weaker heavy-tailedness are assigned larger learning rates to accelerate their training, while layers with stronger heavy-tailedness receive smaller learning rates. By tailoring learning rates in this manner, LLR promotes balanced training across layers, leading to faster convergence and improved generalization. Extensive experiments across architectures (from LLaMA to GPT-nano), optimizers (AdamW and Muon), and parameter scales (60M-1B) demonstrate that LLR achieves up to 1.5x training speedup and outperforms baselines, notably raising average zero-shot accuracy from 47.09% to 49.02%. A key advantage of LLR is its low tuning overhead: it transfers nearly optimal LR settings directly from the uniform baseline. Code is available at https://github.com/hed-ucas/Layer-wise-Learning-Rate.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.