Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
AuthorsJiaqi Deng
AffiliationsIndependent Researcher
Resources
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
Key results
AUC from the joint forward-pass probe on English confirm PAWS-X
AUC from a linear probe over independently encoded sentence vectors
BGE-reranker-large AUC on overlap-matched PAWS-X
PAWS pairs required for nonlinear readers over frozen encodings to recover partial identity
Accuracy after combining cosine retrieval with a joint reader
Top-1 accuracy from cosine alone on meaning-changed wording twins
What the paper found
This paper argues that meaning identity is computed over a sentence pair, not shipped inside independently generated embeddings. On overlap-matched PAWS-X, cosine similarity and linear late fusion of separate vectors from BGE, E5, GTE, MiniLM, Llama 3, Mistral, and Qwen2.5 remain near chance, while probing a mid-depth state from a joint forward pass reaches 0.950 AUC for Qwen2.5-14B versus 0.541 for late fusion; shuffling the partner collapses the signal to chance. The effect generalizes across Qwen2.5, Llama 3, Mistral, GPT-2 XL, DeBERTa, RoBERTa, and Flan-T5, showing it is not a decoder-only or Qwen-specific artifact. BGE-reranker-large reaches 0.940, but Jina reranker-v2 remains at 0.639, so joint processing is necessary rather than sufficient: the model must learn the identity operator. Nonlinear readers over frozen independent encodings recover only part of the relation after 49k PAWS pairs, roughly 40 times the data needed by the joint probe. In an FAQ simulation, adding a joint reader raises trapped-query top-1 accuracy from 0.468 with cosine alone to 0.900 when fused, while preserving the embedding index for cheap retrieval. The practical conclusion is architectural: use embeddings for candidate retrieval, then use a joint model pass to certify whether two texts actually mean the same thing.
Original abstract
Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship. On overlap-matched PAWS-X, purpose-built encoders (BGE, E5, GTE, MiniLM, E5-Mistral-7B) reach English confirm AUC only 0.55-0.65 (dense peak 0.70). Independently encoded last-token states of Llama 3, Mistral, and Qwen do no better; late fusion of the two vectors stays near chance. The same probe on a joint forward pass reaches 0.90-0.96 from 1.5B to 32B, collapses under partner shuffle, is mid-depth, saturates near 0.94 by 3B, and appears more weakly in GPT-2 XL (0.76). The gap holds beyond Llama-style models on other causal LMs, bidirectional encoders (DeBERTa, RoBERTa), and encoder-decoders (Flan-T5, T5, BART). Fixed or linear readers over frozen independent encodings never unlock identity; nonlinear pair readers recover part of it only on the full 49k-pair PAWS train split (0.68-0.87). Off-the-shelf rerankers split: BGE-reranker-large reaches 0.94, while MS-MARCO and Jina stay at 0.55-0.64. Independently trained families compute the same relation and a 1.5B joint reader can distill it from unlabelled teacher scores, while no linear function of the teachers own independent vectors can. Bi-encoders can be fine-tuned to fit PAWS (0.87-0.93), but transfer and STS-B suffer. Cosine compares wording neighbourhoods; identity is a cheap computed operator, not a property of either sentence vector.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
Ej Zhou, Suchir Salhan, Catherine Arnett, Anna Korhonen
Independently trained language models may spontaneously learn representations that can be rotated and transferred across languages, suggesting multilingual capabilities without joint training.