NTH

Variable-Width Transformers

AuthorsZhaofeng Wu, Oliver Sieberling, Shawn Tan, Rameswar Panda, Yury Polyanskiy, Yoon Kim

June 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes making transformers wide at the edges and narrow in the middle, showing that smarter width allocation can improve language model efficiency and performance.

Key results

1B
MoE active parameters

Mixture-of-experts model with 3B total and 1B active parameters

2.751
loss reduction at 2B

constant-width transformer loss at 2B, compared with 2.726 for ><former

16.92
PFLOP/s-days at 2B

constant-width training compute at 2B, compared with 16.49 for ><former

1600
average layer size at 2B

constant-width layer width at 2B, compared with 1426 for ><former

77.8%
FLOPs-needed scaling

FLOPs required by ><former to match the 2B constant-width transformer loss

What the paper found

MIT and MIT-IBM Watson AI Lab researchers introduce Variable-Width Transformers, or “><formers,” a decoder-only architecture that breaks the standard assumption of constant hidden width across depth by using a ×-shaped schedule: wide early layers, a narrow middle bottleneck, and wide late layers. The key implementation is parameter-free residual resizing, where truncated dimensions are carried forward through a fixed global residual stream and restored when width expands, avoiding learned projection bottlenecks. Across dense models from 200M to 2B parameters and a 3B total/1B active MoE model, the variable-width design consistently beats parameter-matched constant-width baselines on language-model loss while also reducing average layer size, which translates to smaller KV cache and lower compute. At the 2B scale, loss drops from 2.751 to 2.726, PFLOP/s-days fall from 16.92 to 16.49, and average layer size falls from 1600 to 1426; scaling-curve fits suggest the same loss can be reached with 77.8% of the FLOPs and 85.1% of the average layer width. On downstream lm-evaluation-harness benchmarks, the 2B model improves average accuracy from 56.1 to 57.2 and LAMBADA perplexity from 8.18 to 7.43, while the MoE variant improves WikiText perplexity from 16.36 to 15.98. Analyses show denser MLP activation use, higher residual-stream entropy in middle layers, and less compression-valley collapse than standard transformers.

Original abstract

Scaling model size, specifically depth and width, has driven significant progress in transformer-based language models. However, most architectures maintain a constant width across all layers, allocating a fixed parameter and computation budget evenly despite different layers potentially playing distinct computational roles. In this work, we empirically investigate nonuniform capacity allocation across network depth by proposing a $\times$-shaped > <former architecture. This design maintains wider early and late layers while narrowing the middle layers, utilizing a parameter-free residual resizing mechanism. Across decoder-only language models ranging from 200M to 2B parameters (dense) and 3B parameters (MoE), our > <former consistently outperforms parameter-matched uniform baselines on language modeling loss. By reducing the average layer width, this architecture also requires fewer overall FLOPs (22% reduction under fitted loss-matched scaling curves) and smaller KV cache memory and I/O cost (15% reduction). In analysis, we show that this bottleneck structure results in qualitatively different representations in residual streams. Overall, our results demonstrate that nonuniform width allocation can result in more resource-optimal scaling of language models.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis