NTH

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

AuthorsAryo Pradipta Gema, Beatrice Alex, Pasquale Minervini

July 4, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes LOCOS, a new way to find attention heads that retrieve meaning rather than just copied words, and shows those heads are crucial for long-context question answering.

Key results

0.401
Qwen3-8B NoLiMa ROUGE-L

baseline before top-50 LOCOS head ablation

0.000
Qwen3-8B NoLiMa ROUGE-L

after ablating the top-50 LOCOS heads

0.292
Wu/NIAH baseline ROUGE-L

strongest attention-based baseline at k=50 on Qwen3-8B

0.55
Qwen3-8B MuSiQue

baseline accuracy before top-50 LOCOS head ablation

0.08
Qwen3-8B MuSiQue

after ablating the top-50 LOCOS heads

0.62
Qwen3-8B BABILong

baseline accuracy before top-50 LOCOS head ablation

What the paper found

This paper, by researchers at the University of Edinburgh, Heriot-Watt University, and Miniml.AI, argues that prior retrieval-head detectors miss a crucial class of heads because they score only where attention reads, not what the output-value circuit writes. The proposed Logit-Contribution Scoring, or LOCOS, projects each head’s OV-circuit output onto the answer-token unembedding direction and applies a needle-versus-off-needle spatial contrast in a single forward pass, making it sensitive to non-literal retrieval on the NoLiMa benchmark. Across six configurations spanning Qwen3-8B, Qwen3-14B, Qwen3-32B, Gemma-3-12B, Gemma-3-27B, and OLMo-3.1-32B, mean-ablation of the top LOCOS heads produces the steepest ROUGE-L collapse; on Qwen3-8B, top-50 ablation drives ROUGE-L from 0.401 to 0.000, while the strongest attention-based baseline still retains 0.292. The selected heads are retrieval-specific rather than generally important: on the same Qwen3-8B model, ablating the top-50 LOCOS heads drops MuSiQue from 0.55 to 0.08 and BABILong from 0.62 to 0.20, while parametric recall and arithmetic remain near baseline. The paper also shows that bottom-k heads with equally large absolute logit contribution but off-needle-dominant signal do not harm retrieval, supporting the causal interpretation. The main novelty is that LOCOS identifies the non-literal subset of retrieval heads that token-matching and attention-mass methods systematically miss.

Original abstract

In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution Scoring (LOCOS), a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass. Across three model families (Qwen3, Gemma-3, OLMo-3.1), mean-ablating the top LOCOS heads on the NoLiMa non-literal retrieval benchmark collapses ROUGE-L at lower head counts than prior attention-based detections; on Qwen3-8B, ablating 50 heads drives ROUGE-L from 0.401 to 0.000 while the strongest baseline still retains 0.292. The selected heads are retrieval-specific: parametric recall and arithmetic reasoning stay at baseline under the same ablation. On Qwen3-8B, the same ablation also drops MuSiQue from 0.55 to 0.08 and BABI-Long from 0.62 to 0.20, while a random-heads control stays within 0.05 of baseline.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis