NTH

Disentangling Representation Evolution in Transformers through Directional Decomposition

AuthorsShwai He, Haichao Zhang, Shen Yan

September 17, 2026 2 min read
Watch on YouTube
The one-line take

This work studies how Transformer layers change representations by separating updates that reinforce existing directions from those that redirect them, revealing implications for editing, compression, and training.

Key results

296M
Pretraining model scale range

Smallest GPT-style model scale evaluated during from-scratch pretraining.

2.7B
Largest pretraining model scale

Largest model scale evaluated for parallel-attention suppression.

1.5
Qwen3-1.7B value-space edit gap

Point gap from the baseline across seven zero-shot benchmarks.

0.1
Qwen3-30B-A3B value-space edit gap

Maximum point gap from baseline when preserving the direct self-value message.

+0.7
1.4B downstream gain

Average accuracy improvement from value-space parallel removal.

What the paper found

This paper introduces directional decomposition for analyzing how Transformer representations evolve: every learned update is separated into a parallel component, which rescales the incoming representation, and a perpendicular component, which redirects it into new semantic directions. Across Qwen3, Llama-3, and Gemma-3 models—including Qwen3 work associated with ByteDance—the experiments show a consistent asymmetry: perpendicular edits sharply damage perplexity and task accuracy, while parallel edits are comparatively redundant, especially when attention preserves the token’s direct self-value message and scales only cross-token aggregation. On seven zero-shot benchmarks, this value-space exclude-self intervention trails Qwen3-1.7B’s baseline by only 1.5 points and stays within 0.1 points on Qwen3-30B-A3B. The same geometry improves diagnostics: under 4-bit AWQ and Wanda pruning, perpendicular compression error tracks downstream fidelity, with correlations of at least 0.97, whereas parallel error is poorly predictive. Finally, suppressing parallel attention components during GPT-style pretraining on OpenWebText improves downstream averages across models from 296M to 2.7B parameters, producing gains of +0.7 points at 1.4B and +1.5 points at 2.7B, with the value-space variant strongest. The result is a unified framework linking representation editing, compression quality, and training-time inductive bias, while challenging isotropic L2 error as a sufficient measure of Transformer distortion.

Original abstract

Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the residual identity path. We then apply the decomposition in two spaces: to attention and MLP updates relative to the hidden state, and to attention value aggregation relative to the current token's value. Targeted edits reveal a strongly space-dependent asymmetry: exclude-self value-space parallel manipulation is markedly more robust than residual-space and perpendicular counterparts, preserving the direct self message while scaling only the non-self aggregate. The same decomposition gives a component-resolved description of compression-induced update error: perpendicular error separates compression methods more clearly than parallel error. Extensive experiments further demonstrate that full-aggregate parallel suppression during from-scratch pretraining lowers validation-loss trajectories and improves downstream averages, with the value-space variant strongest. Together, these results connect representation geometry to editing robustness, compression diagnosis, and training-time intervention. Code is available in the \href{https://github.com/Shwai-He/Transformer-Geometry}{project repository}.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis