Disentangling Representation Evolution in Transformers through Directional Decomposition
AuthorsShwai He, Haichao Zhang, Shen Yan
This work studies how Transformer layers change representations by separating updates that reinforce existing directions from those that redirect them, revealing implications for editing, compression, and training.
Key results
Smallest GPT-style model scale evaluated during from-scratch pretraining.
Largest model scale evaluated for parallel-attention suppression.
Point gap from the baseline across seven zero-shot benchmarks.
Maximum point gap from baseline when preserving the direct self-value message.
Average accuracy improvement from value-space parallel removal.
What the paper found
This paper introduces directional decomposition for analyzing how Transformer representations evolve: every learned update is separated into a parallel component, which rescales the incoming representation, and a perpendicular component, which redirects it into new semantic directions. Across Qwen3, Llama-3, and Gemma-3 models—including Qwen3 work associated with ByteDance—the experiments show a consistent asymmetry: perpendicular edits sharply damage perplexity and task accuracy, while parallel edits are comparatively redundant, especially when attention preserves the token’s direct self-value message and scales only cross-token aggregation. On seven zero-shot benchmarks, this value-space exclude-self intervention trails Qwen3-1.7B’s baseline by only 1.5 points and stays within 0.1 points on Qwen3-30B-A3B. The same geometry improves diagnostics: under 4-bit AWQ and Wanda pruning, perpendicular compression error tracks downstream fidelity, with correlations of at least 0.97, whereas parallel error is poorly predictive. Finally, suppressing parallel attention components during GPT-style pretraining on OpenWebText improves downstream averages across models from 296M to 2.7B parameters, producing gains of +0.7 points at 1.4B and +1.5 points at 2.7B, with the value-space variant strongest. The result is a unified framework linking representation editing, compression quality, and training-time inductive bias, while challenging isotropic L2 error as a sufficient measure of Transformer distortion.
Original abstract
Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the residual identity path. We then apply the decomposition in two spaces: to attention and MLP updates relative to the hidden state, and to attention value aggregation relative to the current token's value. Targeted edits reveal a strongly space-dependent asymmetry: exclude-self value-space parallel manipulation is markedly more robust than residual-space and perpendicular counterparts, preserving the direct self message while scaling only the non-self aggregate. The same decomposition gives a component-resolved description of compression-induced update error: perpendicular error separates compression methods more clearly than parallel error. Extensive experiments further demonstrate that full-aggregate parallel suppression during from-scratch pretraining lowers validation-loss trajectories and improves downstream averages, with the value-space variant strongest. Together, these results connect representation geometry to editing robustness, compression diagnosis, and training-time intervention. Code is available in the \href{https://github.com/Shwai-He/Transformer-Geometry}{project repository}.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.