NTH

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

AuthorsChanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung

September 3, 2026 2 min read
Watch on YouTube
The one-line take

This paper finds a tiny set of attention heads that act like visual search engines inside vision-language models, locating the image evidence needed for their answers.

Key results

11
Evaluated VLMs

Number of VLMs tested for VRH universality.

5
Grounding benchmarks

Number of referring-expression grounding benchmarks used.

20
Masked VRHs

Top-ranked heads masked during causal evaluation.

2.6%
VRH sparsity upper bound

Largest reported fraction of model attention heads represented by the top 20 VRHs.

80
Maximum grounding drop

Maximum reduction in grounding accuracy measured in percentage points after VRH masking.

0.78
DPO improvement

Average grounding improvement in percentage points from VRH-targeted Direct Preference Optimization.

What the paper found

This paper introduces Visual Retrieval Heads, or VRHs: a sparse subset of attention heads that causally routes the visual evidence needed to resolve a text reference. The researchers identify them by scoring attention from output prediction tokens to visual tokens inside the ground-truth referent, then validate the result by masking those heads during inference. Across 11 VLMs and 5 grounding benchmarks, including Qwen2.5-VL, Qwen3-VL, InternVL, Llama-3-based systems, and DeepSeek-VL2, masking the top 20 VRHs—only 1.7–2.6% of all heads—reduces grounding accuracy by up to 80 percentage points, while masking random heads has little effect. The same heads also remain causal for attributes, spatial reasoning, counting, and visual mathematics on benchmarks such as VAW, Spatial457, CountBenchQA, and MathVista. VRHs preserve output syntax while causing spatially incorrect boxes or fluent but visually unfaithful answers, showing functional specificity rather than general generation damage. They transfer across VLMs that share an LLM backbone despite different vision encoders, projectors, and instruction tuning, and can be detected from as few as 5–10 annotated examples instead of 200. As an intervention, Direct Preference Optimization using VRH-masked negatives improves grounding by an average of 0.78 percentage points, suggesting these heads can support both interpretability and targeted model optimization.

Original abstract

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under a unified design space over query tokens, key aggregation, and cross-sample aggregation. We then show that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads. Across eleven VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect. Beyond replicating the causal-sparse-universal triad established for text retrieval heads, VRHs exhibit several properties not previously reported: they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction; they are functionally specific, preserving output format while corrupting localization; and they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis