NTH

Tapered Language Models

AuthorsReza Bayat, Ali Behrouz, Aaron Courville

July 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that language models should spend more capacity in earlier layers and less in later ones, and shows that a tapered width schedule can improve performance without increasing cost.

Key results

16.28
440M baseline perplexity

Uniform 440M Transformer validation perplexity

14.44
Cosine taper perplexity

Best cosine-tapered 440M Transformer validation perplexity

17.29
Wider-late perplexity

Reverse allocation on 440M Transformer

52.25
760M Transformer++ avg accuracy

Uniform baseline average commonsense accuracy

52.84
760M Transformer++ tapered avg accuracy

Tapered average commonsense accuracy

56.05
1.3B Transformer++ avg accuracy

Uniform baseline average commonsense accuracy

What the paper found

Tapered Language Models, from Mila and Cornell University with Aaron Courville, challenge the default assumption that every layer in a language model should get the same parameter budget. The paper shows that when MLP width is redistributed across depth under a fixed parameter and FLOP budget, front-loading capacity consistently beats uniform allocation: on a 440M Transformer, a cosine taper reaches 14.44 validation perplexity versus 16.28 for the uniform baseline, while the reverse, back-loaded allocation rises to 17.29. Using the best cosine schedule with a 1.5/0.5 start-to-end width ratio, the authors report gains across four architectures—Transformer, Gated Attention, Hope-attention, and Titans—at 760M and 1.3B parameters: for example, average commonsense accuracy improves from 52.25 to 52.84 on Transformer++ at 760M and from 56.05 to 56.38 at 1.3B, with perplexity improving on WikiText and LAMBADA in nearly all comparisons. The mechanism is supported by layer-wise cosine-similarity analysis on pretrained GPT-2 checkpoints from 124M to 1.5B parameters, where both full-block and MLP updates become more aligned with the residual stream as depth increases, implying that later MLPs write less novel information. The key technical takeaway is that depth-aware capacity allocation is a free architectural lever: tapering MLP width improves language modeling, commonsense reasoning, and long-context retrieval without increasing parameters or compute.

Original abstract

Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This is a default inherited from the original transformer and largely unchanged since, yet a growing body of evidence suggests that layers contribute non-uniformly to the final output, with later layers refining the residual stream rather than transforming it. We ask whether parameter capacity should reflect this asymmetry. Our controlled experiment shows that, under a fixed budget, allocating more capacity to earlier layers and less to later layers improves perplexity over a uniform-width baseline, while the reverse allocation hurts. Building on this result, we introduce Tapered Language Models (TLMs), an architectural principle in which a parameter-bearing component is monotonically tapered across depth under a fixed total budget. MLPs are the natural site for this instantiation: they dominate parameter count across all modern LM families and expose width as a single, clean axis of variation. Across three model scales and four architectures (Transformer, Gated Attention, Hope-attention, and Titans), tapering MLP width via a smooth cosine schedule consistently improves perplexity and downstream benchmark performance over uniform baselines, at no additional parameter or compute cost. These findings establish depth-aware capacity allocation as a simple, architecture-agnostic axis of language model design, a free lever hidden in plain sight.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis