Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
AuthorsAryo Pradipta Gema, Beatrice Alex, Pasquale Minervini
Resources
This paper proposes LOCOS, a new way to find attention heads that retrieve meaning rather than just copied words, and shows those heads are crucial for long-context question answering.
Key results
baseline before top-50 LOCOS head ablation
after ablating the top-50 LOCOS heads
strongest attention-based baseline at k=50 on Qwen3-8B
baseline accuracy before top-50 LOCOS head ablation
after ablating the top-50 LOCOS heads
baseline accuracy before top-50 LOCOS head ablation
What the paper found
This paper, by researchers at the University of Edinburgh, Heriot-Watt University, and Miniml.AI, argues that prior retrieval-head detectors miss a crucial class of heads because they score only where attention reads, not what the output-value circuit writes. The proposed Logit-Contribution Scoring, or LOCOS, projects each head’s OV-circuit output onto the answer-token unembedding direction and applies a needle-versus-off-needle spatial contrast in a single forward pass, making it sensitive to non-literal retrieval on the NoLiMa benchmark. Across six configurations spanning Qwen3-8B, Qwen3-14B, Qwen3-32B, Gemma-3-12B, Gemma-3-27B, and OLMo-3.1-32B, mean-ablation of the top LOCOS heads produces the steepest ROUGE-L collapse; on Qwen3-8B, top-50 ablation drives ROUGE-L from 0.401 to 0.000, while the strongest attention-based baseline still retains 0.292. The selected heads are retrieval-specific rather than generally important: on the same Qwen3-8B model, ablating the top-50 LOCOS heads drops MuSiQue from 0.55 to 0.08 and BABILong from 0.62 to 0.20, while parametric recall and arithmetic remain near baseline. The paper also shows that bottom-k heads with equally large absolute logit contribution but off-needle-dominant signal do not harm retrieval, supporting the causal interpretation. The main novelty is that LOCOS identifies the non-literal subset of retrieval heads that token-matching and attention-mass methods systematically miss.
Original abstract
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution Scoring (LOCOS), a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass. Across three model families (Qwen3, Gemma-3, OLMo-3.1), mean-ablating the top LOCOS heads on the NoLiMa non-literal retrieval benchmark collapses ROUGE-L at lower head counts than prior attention-based detections; on Qwen3-8B, ablating 50 heads drives ROUGE-L from 0.401 to 0.000 while the strongest baseline still retains 0.292. The selected heads are retrieval-specific: parametric recall and arithmetic reasoning stay at baseline under the same ablation. On Qwen3-8B, the same ablation also drops MuSiQue from 0.55 to 0.08 and BABI-Long from 0.62 to 0.20, while a random-heads control stays within 0.05 of baseline.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.