NTH

The Geometry of Inference in Transformer Residual Streams

AuthorsTimur Mudarisov, Mikhail Burtsev, Radu State

AffiliationsUniversity of Luxembourg, Luxembourg · London Institute for Mathematical Sciences, London, UK

October 8, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.

Key results

6
Models evaluated

Pretrained Gemma, Qwen2.5, Mistral-7B, and Llama-3 models.

1024
Contexts per model

Non-overlapping contexts sampled from FineWeb sample-10BT.

256
Context length

Tokens in each analyzed context.

1023
Alternative endpoints per query

Primary endpoint bank excluding the query’s own final state.

7.5
Endpoint diagnostic improvement

Maximum improvement factor over the uniform-sphere baseline.

5
Projected-normal best-model count

Models in which projected normal gave the lowest synthetic cosine prediction error.

What the paper found

This paper proposes viewing Transformer inference as geometric refinement in the residual stream: each intermediate state is compared with its own final residual endpoint and with endpoints from other contexts, called competitors when they are closer. Across 6 pretrained models—Google’s Gemma-2B and Gemma-7B, Qwen2.5-1.5B and Qwen2.5-7B, Mistral-7B, and Meta’s Llama-3-8B—using 1024 non-overlapping 256-token contexts from FineWeb sample-10BT, the own endpoint is closer than the average alternative from the first measured layer, even though as many as hundreds of individual competitors remain; the primary bank provides 1023 alternatives per query. Competitor counts generally decline with depth, but cosine-distance sets show both entries and exits, proving that inference revises endpoint preferences rather than simply eliminating alternatives along a straight path. The analysis separates norm, directional alignment, and endpoint-population geometry, explaining how alignment can improve while Euclidean distance stays nearly flat and why small directional changes can sharply reduce competition in high dimensions. Final endpoints linked to lower-ranked output tokens are consistently farther away in cosine distance. To model endpoint populations, the study fits spherical-cap, von Mises–Fisher, mixture, and projected-normal distributions; structured fits improve endpoint diagnostics by as much as 7.5-fold over a uniform-sphere baseline, while the projected-normal model predicts held-out cosine competitor fractions most accurately in 5 of 6 models. The result is a quantitative account of how residual updates develop increasing specificity without implying an explicit internal search process.

Original abstract

Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.

Read the original paper

More in Transformers

Browse all 42 papers →
02Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis