The Geometry of Inference in Transformer Residual Streams
AuthorsTimur Mudarisov, Mikhail Burtsev, Radu State
AffiliationsUniversity of Luxembourg, Luxembourg · London Institute for Mathematical Sciences, London, UK
Resources
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Key results
Pretrained Gemma, Qwen2.5, Mistral-7B, and Llama-3 models.
Non-overlapping contexts sampled from FineWeb sample-10BT.
Tokens in each analyzed context.
Primary endpoint bank excluding the query’s own final state.
Maximum improvement factor over the uniform-sphere baseline.
Models in which projected normal gave the lowest synthetic cosine prediction error.
What the paper found
This paper proposes viewing Transformer inference as geometric refinement in the residual stream: each intermediate state is compared with its own final residual endpoint and with endpoints from other contexts, called competitors when they are closer. Across 6 pretrained models—Google’s Gemma-2B and Gemma-7B, Qwen2.5-1.5B and Qwen2.5-7B, Mistral-7B, and Meta’s Llama-3-8B—using 1024 non-overlapping 256-token contexts from FineWeb sample-10BT, the own endpoint is closer than the average alternative from the first measured layer, even though as many as hundreds of individual competitors remain; the primary bank provides 1023 alternatives per query. Competitor counts generally decline with depth, but cosine-distance sets show both entries and exits, proving that inference revises endpoint preferences rather than simply eliminating alternatives along a straight path. The analysis separates norm, directional alignment, and endpoint-population geometry, explaining how alignment can improve while Euclidean distance stays nearly flat and why small directional changes can sharply reduce competition in high dimensions. Final endpoints linked to lower-ranked output tokens are consistently farther away in cosine distance. To model endpoint populations, the study fits spherical-cap, von Mises–Fisher, mixture, and projected-normal distributions; structured fits improve endpoint diagnostics by as much as 7.5-fold over a uniform-sphere baseline, while the projected-normal model predicts held-out cosine competitor fractions most accurately in 5 of 6 models. The result is a quantitative account of how residual updates develop increasing specificity without implying an explicit internal search process.
Original abstract
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.
Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim
This work shows, in theory, how transformers can adapt to data living on locally different geometric structures and still achieve statistically optimal in-context prediction.