Resources
This paper makes transformer feedforward layers both smaller and more interpretable by replacing dense expert blocks with sparsely selected single-neuron linear experts.
Key results
Language-model training and isoflop comparisons were run on the SlimPajama corpus with 627B tokens.
The isoflop comparison evaluated four FLOP budgets spanning 1×10^17 to 6×10^18 FLOPs.
Circuit analysis collected gating weights from a random subset of 128 unseen TinyStories validation stories.
What the paper found
Simon Schug’s “Sparsely gated tiny linear experts” proposes sgatlin, a transformer feedforward layer that pushes sparsity to an extreme: each expert is a single linear neuron, and only k=8 neurons are activated per token out of a very large pool, with a small gating bottleneck of d_key=128. The key design choice is removing the usual expert nonlinearity entirely, so the selected feedforward circuit is, conditioned on the gate, a low-rank linear map. In compute-matched isoflop training on the 627B-token SlimPajama corpus, replacing every transformer MLP with sgatlin improves test perplexity across budgets from 1×10^17 to 6×10^18 FLOPs and remains competitive against dense GeLU and SwiGLU layers as well as MoE baselines, including OpenAI’s GPT-OSS-style coarse MoE and the fine-grained PEER architecture from Xu Owen He. Ablations show that adding ReLU, GeLU, or Swish worsens perplexity, confirming that hard top-k gating alone is sufficient. The paper’s interpretability result is especially novel: on a TinyStories model, gating weights form semantically structured clusters under UMAP, nearest-neighbor circuit search retrieves related contexts such as names, pronouns, and animal categories, and causal patching of gating weights measurably shifts factual recall, with the strongest normalized indirect effects appearing when intervening at noun positions and especially in the second layer.
Original abstract
Sparsity allows scaling model parameters without proportionally increasing computational cost. While mixture of experts (MoE) models are made increasingly sparse, individual experts typically remain large and dense. Here, we demonstrate that further increasing sparsity by shrinking each expert to consist of a single neuron and selecting a tiny fraction of many available neurons can improve compute efficiency and interpretability. Counterintuitively, the key to achieving both is removing the nonlinearity typically applied to the experts, resulting in a network of sparsely gated linear neurons (sgatlin). In an isoflop comparison, we find that replacing all transformer feedforward layers with sgatlin improves perplexity in language models across different compute budgets. At the same time, the sparsity and linearity of the resulting feedforward circuits present new opportunities for model interpretability. In a small-scale case study, we demonstrate that feedforward circuits in sgatlin can be interpreted without having to train additional replacement models. We find that they form semantically structured clusters and are causally implicated in factual recall. Our findings paint a possible path towards compute-efficient and interpretable transformer feedforward layers.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.