Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
AuthorsPierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
Resources
This paper separates what Transformers know about meaning and word order into distinct streams, revealing how positional structure is stored and showing that the split can improve representations.
Key results
The learned absolute positional embedding matrix in DSTG-NeoBERT collapses into a two-dimensional low-frequency manifold, with the first two principal components explaining 94.1% of the variance.
The model is trained on the English FineWeb corpus for about 22 billion tokens using masked language modeling.
The disentangled model improves results on 49 of the 65 linguistic phenomena in the Flash-Holmes probing benchmark.
What the paper found
In “Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders,” Pierre-Antoine Lequeu, Camille Barboule from Orange Innovation, and Benjamin Piwowarski at Sorbonne Université and CNRS propose a Transformer encoder that separates token content into three streams: semantic, absolute positional, and relative positional information. Built on NeoBERT and trained on 22 billion FineWeb tokens with masked language modeling, the key change is that the MLM head predicts only from the semantic stream, leaving the absolute-position stream free of prediction pressure. Mechanistically, this reveals that learned absolute positions collapse into a low-dimensional, two-dimensional low-frequency manifold capturing document structure, with 94.1% of variance in the first two principal components, while attention heads specialize into structure-oriented and semantic-oriented groups; relative-position bias supports semantic retrieval but does not become a standalone positional representation. Standard encoders with RoPE or relative bias weakly retain macroscopic structure, and entangled learned absolute embeddings lose it in later layers. On probing, the disentangled model matches or exceeds baselines on GLUE, MTEB, and SQuAD, and improves 49 of 65 Flash-Holmes linguistic phenomena, especially syntax, semantics, and discourse, showing that explicitly preserving positional structure can yield richer token representations without harming mainstream performance.
Original abstract
Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks such as long-context understanding or retrieval \cite{chen-etal-2025-hope}. Hence, a better understanding of the internal positional mechanism could help design better PE. Building on evidence that positional and semantic signals occupy nearly orthogonal subspaces in trained Transformers, we modify an encoder Transformer to process three explicitly disentangled streams: semantic, absolute positional (AP) and relative positional (RP), and confine the masked-language-modeling (MLM) objective to the semantic stream. This decoupling enables a clean mechanistic study and yields three take-aways. (1) The isolated AP subspace spontaneously collapses into a low-frequency two-dimensional manifold that captures the structure of the document; (2) Attention heads specialize into structure and semantic-oriented groups, with RP exclusively supporting the latter; (3) Standard positional encodings do not robustly retain macroscopic structure: RoPE and RP only weakly encode it, and entangled AP loses it in the final layers under MLM pressure. The disentangled approach preserves positional encoding, which improves linguistic representation on 49 of the 65 linguistic phenomena of the Flash-Holmes probing benchmark.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.