NTH

Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings

AuthorsJiaqi Deng

AffiliationsIndependent Researcher

October 1, 2026 3 min read
Watch on YouTube
The one-line take

The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.

Key results

0.950
Qwen2.5-14B joint AUC

AUC from the joint forward-pass probe on English confirm PAWS-X

0.541
Qwen2.5-14B late-fusion AUC

AUC from a linear probe over independently encoded sentence vectors

0.940
Reranker comparison

BGE-reranker-large AUC on overlap-matched PAWS-X

49k
Independent-reader training scale

PAWS pairs required for nonlinear readers over frozen encodings to recover partial identity

0.900
FAQ fused top-1 accuracy

Accuracy after combining cosine retrieval with a joint reader

0.468
FAQ cosine trap accuracy

Top-1 accuracy from cosine alone on meaning-changed wording twins

What the paper found

This paper argues that meaning identity is computed over a sentence pair, not shipped inside independently generated embeddings. On overlap-matched PAWS-X, cosine similarity and linear late fusion of separate vectors from BGE, E5, GTE, MiniLM, Llama 3, Mistral, and Qwen2.5 remain near chance, while probing a mid-depth state from a joint forward pass reaches 0.950 AUC for Qwen2.5-14B versus 0.541 for late fusion; shuffling the partner collapses the signal to chance. The effect generalizes across Qwen2.5, Llama 3, Mistral, GPT-2 XL, DeBERTa, RoBERTa, and Flan-T5, showing it is not a decoder-only or Qwen-specific artifact. BGE-reranker-large reaches 0.940, but Jina reranker-v2 remains at 0.639, so joint processing is necessary rather than sufficient: the model must learn the identity operator. Nonlinear readers over frozen independent encodings recover only part of the relation after 49k PAWS pairs, roughly 40 times the data needed by the joint probe. In an FAQ simulation, adding a joint reader raises trapped-query top-1 accuracy from 0.468 with cosine alone to 0.900 when fused, while preserving the embedding index for cheap retrieval. The practical conclusion is architectural: use embeddings for candidate retrieval, then use a joint model pass to certify whether two texts actually mean the same thing.

Original abstract

Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship. On overlap-matched PAWS-X, purpose-built encoders (BGE, E5, GTE, MiniLM, E5-Mistral-7B) reach English confirm AUC only 0.55-0.65 (dense peak 0.70). Independently encoded last-token states of Llama 3, Mistral, and Qwen do no better; late fusion of the two vectors stays near chance. The same probe on a joint forward pass reaches 0.90-0.96 from 1.5B to 32B, collapses under partner shuffle, is mid-depth, saturates near 0.94 by 3B, and appears more weakly in GPT-2 XL (0.76). The gap holds beyond Llama-style models on other causal LMs, bidirectional encoders (DeBERTa, RoBERTa), and encoder-decoders (Flan-T5, T5, BART). Fixed or linear readers over frozen independent encodings never unlock identity; nonlinear pair readers recover part of it only on the full 49k-pair PAWS train split (0.68-0.87). Off-the-shelf rerankers split: BGE-reranker-large reaches 0.94, while MS-MARCO and Jina stay at 0.55-0.64. Independently trained families compute the same relation and a 1.5B joint reader can distill it from unlabelled teacher scores, while no linear function of the teachers own independent vectors can. Bi-encoders can be fine-tuned to fit PAWS (0.87-0.93), but transfer and STS-B suffer. Cosine compares wording neighbourhoods; identity is a cheap computed operator, not a property of either sentence vector.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →