NTH

Xetrieval: Mechanistically Explaining Dense Retrieval

AuthorsZhixin Cai, Jun Bai, Yang Liu, Jiaqi Li, Yichi Zhang, Taichuan Li, Zhuofan Chen, Zixia Jia, Zilong Zheng, Wenge Rong

June 10, 2026 2 min read
Watch on YouTube
The one-line take

Xetrieval explains why dense retrievers rank documents highly by turning opaque embeddings into sparse human-readable features and using them to trace retrieval decisions.

Key results

11796
StackExchange corpus

documents used to train the reasoning internalizer

256
TopK-SAE k

chosen backbone sparsity level for the mechanistic explainer

61.5
e5-large avg NDCG@10 baseline

unenhanced dense retriever average on 7 benchmarks

64.2
e5-large avg NDCG@10 with reasoning internalizer

average retrieval score after embedding-level reasoning enhancement

What the paper found

Xetrieval, developed by researchers from Beihang University and BIGAI, explains dense retrieval by decomposing query and document embeddings into sparse, human-readable features instead of relying on lexical rationales. The framework has two parts: a reasoning internalizer, trained on 11,796 StackExchange documents with LLM-generated Summary, Purpose, and QA reasoning targets, and a mechanistic explainer based on TopK sparse autoencoders with k = 256. On 7 benchmarks including BRIGHT, NQ, MuTual, TREC-NEWS, Signal-1M, Robust04, and ArguAna, the reasoning internalizer improves retrieval over the base retriever and reaches an average NDCG@10 of 64.2 on e5-large versus 61.5 without enhancement, while an explicit CoT reasoner reaches 66.5. In analysis, TopK-SAE at L0 = 256 gives the best balance of reconstruction, mono-semanticity, and retrieval retention. Compared with a raw SAE, the reasoning-augmented explainer produces more coherent features and stronger intervention effects: erasing Xetrieval-selected features causes the largest similarity drop, and amplifying key features improves NDCG@10 on BRIGHT, ArguAna, and NQ more reliably than steering raw SAE features. The main claim is that dense retrieval decisions can be traced to latent query-document factors that are both interpretable and causally relevant.

Original abstract

Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanations often focus on surface signals, such as lexical matches, token alignments, or post-hoc textual rationales, and thus provide limited insight into the latent factors that shape dense retrieval behavior at the embedding level. We propose \textit{Xetrieval}, an embedding-level mechanistic framework for explaining dense retrieval. \textit{Xetrieval} first introduces a lightweight reasoning internalizer that approximates Chain-of-Thought reasoning directly in the embedding space with a single forward pass, enriching sentence embeddings with reasoning-oriented information while avoiding expensive autoregressive generation. It then decomposes these reasoning-enhanced embeddings into sparse, human-interpretable features, each associated with a coherent natural language description. By aggregating sparse feature overlaps across multiple document-side views, \textit{Xetrieval} provides feature-level explanations of individual retrieval decisions. Experiments on diverse retrievers and benchmarks show that \textit{Xetrieval} uncovers coherent interpretable features, yields stronger pair-level intervention effects, and supports task-level feature steering. The project page and source code are available at https://hihiczx.github.io/Xetrieval .

Read the original paper

More in Natural Language Processing

Browse all 26 papers →