Tapered Language Models
AuthorsReza Bayat, Ali Behrouz, Aaron Courville
Resources
This paper argues that language models should spend more capacity in earlier layers and less in later ones, and shows that a tapered width schedule can improve performance without increasing cost.
Key results
Uniform 440M Transformer validation perplexity
Best cosine-tapered 440M Transformer validation perplexity
Reverse allocation on 440M Transformer
Uniform baseline average commonsense accuracy
Tapered average commonsense accuracy
Uniform baseline average commonsense accuracy
What the paper found
Tapered Language Models, from Mila and Cornell University with Aaron Courville, challenge the default assumption that every layer in a language model should get the same parameter budget. The paper shows that when MLP width is redistributed across depth under a fixed parameter and FLOP budget, front-loading capacity consistently beats uniform allocation: on a 440M Transformer, a cosine taper reaches 14.44 validation perplexity versus 16.28 for the uniform baseline, while the reverse, back-loaded allocation rises to 17.29. Using the best cosine schedule with a 1.5/0.5 start-to-end width ratio, the authors report gains across four architectures—Transformer, Gated Attention, Hope-attention, and Titans—at 760M and 1.3B parameters: for example, average commonsense accuracy improves from 52.25 to 52.84 on Transformer++ at 760M and from 56.05 to 56.38 at 1.3B, with perplexity improving on WikiText and LAMBADA in nearly all comparisons. The mechanism is supported by layer-wise cosine-similarity analysis on pretrained GPT-2 checkpoints from 124M to 1.5B parameters, where both full-block and MLP updates become more aligned with the residual stream as depth increases, implying that later MLPs write less novel information. The key technical takeaway is that depth-aware capacity allocation is a free architectural lever: tapering MLP width improves language modeling, commonsense reasoning, and long-context retrieval without increasing parameters or compute.
Original abstract
Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This is a default inherited from the original transformer and largely unchanged since, yet a growing body of evidence suggests that layers contribute non-uniformly to the final output, with later layers refining the residual stream rather than transforming it. We ask whether parameter capacity should reflect this asymmetry. Our controlled experiment shows that, under a fixed budget, allocating more capacity to earlier layers and less to later layers improves perplexity over a uniform-width baseline, while the reverse allocation hurts. Building on this result, we introduce Tapered Language Models (TLMs), an architectural principle in which a parameter-bearing component is monotonically tapered across depth under a fixed total budget. MLPs are the natural site for this instantiation: they dominate parameter count across all modern LM families and expose width as a single, clean axis of variation. Across three model scales and four architectures (Transformer, Gated Attention, Hope-attention, and Titans), tapering MLP width via a smooth cosine schedule consistently improves perplexity and downstream benchmark performance over uniform baselines, at no additional parameter or compute cost. These findings establish depth-aware capacity allocation as a simple, architecture-agnostic axis of language model design, a free lever hidden in plain sight.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.