Variable-Width Transformers
AuthorsZhaofeng Wu, Oliver Sieberling, Shawn Tan, Rameswar Panda, Yury Polyanskiy, Yoon Kim
Resources
This paper proposes making transformers wide at the edges and narrow in the middle, showing that smarter width allocation can improve language model efficiency and performance.
Key results
Mixture-of-experts model with 3B total and 1B active parameters
constant-width transformer loss at 2B, compared with 2.726 for ><former
constant-width training compute at 2B, compared with 16.49 for ><former
constant-width layer width at 2B, compared with 1426 for ><former
FLOPs required by ><former to match the 2B constant-width transformer loss
What the paper found
MIT and MIT-IBM Watson AI Lab researchers introduce Variable-Width Transformers, or “><formers,” a decoder-only architecture that breaks the standard assumption of constant hidden width across depth by using a ×-shaped schedule: wide early layers, a narrow middle bottleneck, and wide late layers. The key implementation is parameter-free residual resizing, where truncated dimensions are carried forward through a fixed global residual stream and restored when width expands, avoiding learned projection bottlenecks. Across dense models from 200M to 2B parameters and a 3B total/1B active MoE model, the variable-width design consistently beats parameter-matched constant-width baselines on language-model loss while also reducing average layer size, which translates to smaller KV cache and lower compute. At the 2B scale, loss drops from 2.751 to 2.726, PFLOP/s-days fall from 16.92 to 16.49, and average layer size falls from 1600 to 1426; scaling-curve fits suggest the same loss can be reached with 77.8% of the FLOPs and 85.1% of the average layer width. On downstream lm-evaluation-harness benchmarks, the 2B model improves average accuracy from 56.1 to 57.2 and LAMBADA perplexity from 8.18 to 7.43, while the MoE variant improves WikiText perplexity from 16.36 to 15.98. Analyses show denser MLP activation use, higher residual-stream entropy in middle layers, and less compression-valley collapse than standard transformers.
Original abstract
Scaling model size, specifically depth and width, has driven significant progress in transformer-based language models. However, most architectures maintain a constant width across all layers, allocating a fixed parameter and computation budget evenly despite different layers potentially playing distinct computational roles. In this work, we empirically investigate nonuniform capacity allocation across network depth by proposing a $\times$-shaped > <former architecture. This design maintains wider early and late layers while narrowing the middle layers, utilizing a parameter-free residual resizing mechanism. Across decoder-only language models ranging from 200M to 2B parameters (dense) and 3B parameters (MoE), our > <former consistently outperforms parameter-matched uniform baselines on language modeling loss. By reducing the average layer width, this architecture also requires fewer overall FLOPs (22% reduction under fitted loss-matched scaling curves) and smaller KV cache memory and I/O cost (15% reduction). In analysis, we show that this bottleneck structure results in qualitatively different representations in residual streams. Overall, our results demonstrate that nonuniform width allocation can result in more resource-optimal scaling of language models.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.