NTH

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

AuthorsDi He, Songjun Tu, Keyu Wang, Lu Yin, Shiwei Liu

May 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper uses heavy-tail theory to automatically give different Transformer layers different learning rates, helping large language models train faster and perform better.

Key results

21.94 to 20.30
FineWeb LLaMA-60M perplexity

On FineWeb pretraining, LLR lowers validation perplexity versus tuned Uniform AdamW for the 60M LLaMA model.

17.86 to 17.03
FineWeb LLaMA-135M perplexity

LLR improves validation perplexity over Uniform AdamW on the 135M LLaMA model during FineWeb pretraining.

12.96 to 12.71
FineWeb LLaMA-350M perplexity

LLR reduces validation perplexity relative to Uniform AdamW for the 350M LLaMA model.

9.77 to 9.59
FineWeb LLaMA-1B perplexity

LLR achieves lower validation perplexity than Uniform AdamW on the 1B LLaMA model.

47.09% to 49.02%
LLaMA-1B zero-shot average accuracy

After pretraining, LLR improves average zero-shot commonsense accuracy across PIQA, SIQA, HellaSwag, WinoGrande, ARC-c, ARC-e, and OBQA.

What the paper found

“One LR Doesn’t Fit All” argues that a single global learning rate is a poor match for Transformer heterogeneity in LLM pretraining, and introduces Layerwise Learning Rate, or LLR, which uses Heavy-Tailed Self-Regularization theory to measure each layer’s empirical spectral density, fit a power law, and compute a PL_Alpha_Hill exponent as a proxy for how “heavy-tailed” and well-trained that layer already is. The rule is simple but novel: layers with lower PL_Alpha_Hill, such as attention projections like Att.q and Att.k, receive smaller learning rates, while layers with higher PL_Alpha_Hill, especially FFN and embedding blocks, receive larger ones; LLaMA embeddings are pinned to the upper LR bound, and LR changes are applied with a soft-switch schedule during only the first 20% of training to avoid spikes and reduce overhead. On FineWeb pretraining, LLR improves validation perplexity over tuned Uniform AdamW from 21.94 to 20.30 on LLaMA-60M, 17.86 to 17.03 on 135M, 12.96 to 12.71 on 350M, and 9.77 to 9.59 on 1B, while also outperforming LARS, LAMB, and the sharpness-based baseline. On LLaMA-1B, zero-shot commonsense accuracy across PIQA, SIQA, HellaSwag, WinoGrande, ARC-c, ARC-e, and OBQA rises from 47.09% to 49.02%. The method transfers to GPT-nano, RoBERTa-base fine-tuning, and Muon optimization, and the authors report up to 1.5× faster convergence with substantially lower tuning overhead than prior layerwise schemes.

Original abstract

Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural heterogeneity of Transformers, potentially limiting their effectiveness as the backbone of Large Language Models (LLMs). In this paper, we introduce Layerwise Learning Rate (LLR), an adaptive scheme that assigns distinct learning rates to individual Transformer layers. Our method is grounded in Heavy-Tailed Self-Regularization (HT-SR) theory, which characterizes the empirical spectral density (ESD) of weight correlation matrices to quantify heavy-tailedness. Layers with weaker heavy-tailedness are assigned larger learning rates to accelerate their training, while layers with stronger heavy-tailedness receive smaller learning rates. By tailoring learning rates in this manner, LLR promotes balanced training across layers, leading to faster convergence and improved generalization. Extensive experiments across architectures (from LLaMA to GPT-nano), optimizers (AdamW and Muon), and parameter scales (60M-1B) demonstrate that LLR achieves up to 1.5x training speedup and outperforms baselines, notably raising average zero-shot accuracy from 47.09% to 49.02%. A key advantage of LLR is its low tuning overhead: it transfers nearly optimal LR settings directly from the uniform baseline. Code is available at https://github.com/hed-ucas/Layer-wise-Learning-Rate.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis