Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information
AuthorsChanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
Resources
This paper finds a tiny set of attention heads that act like visual search engines inside vision-language models, locating the image evidence needed for their answers.
Key results
Number of VLMs tested for VRH universality.
Number of referring-expression grounding benchmarks used.
Top-ranked heads masked during causal evaluation.
Largest reported fraction of model attention heads represented by the top 20 VRHs.
Maximum reduction in grounding accuracy measured in percentage points after VRH masking.
Average grounding improvement in percentage points from VRH-targeted Direct Preference Optimization.
What the paper found
This paper introduces Visual Retrieval Heads, or VRHs: a sparse subset of attention heads that causally routes the visual evidence needed to resolve a text reference. The researchers identify them by scoring attention from output prediction tokens to visual tokens inside the ground-truth referent, then validate the result by masking those heads during inference. Across 11 VLMs and 5 grounding benchmarks, including Qwen2.5-VL, Qwen3-VL, InternVL, Llama-3-based systems, and DeepSeek-VL2, masking the top 20 VRHs—only 1.7–2.6% of all heads—reduces grounding accuracy by up to 80 percentage points, while masking random heads has little effect. The same heads also remain causal for attributes, spatial reasoning, counting, and visual mathematics on benchmarks such as VAW, Spatial457, CountBenchQA, and MathVista. VRHs preserve output syntax while causing spatially incorrect boxes or fluent but visually unfaithful answers, showing functional specificity rather than general generation damage. They transfer across VLMs that share an LLM backbone despite different vision encoders, projectors, and instruction tuning, and can be detected from as few as 5–10 annotated examples instead of 200. As an intervention, Direct Preference Optimization using VRH-masked negatives improves grounding by an average of 0.78 percentage points, suggesting these heads can support both interpretability and targeted model optimization.
Original abstract
Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under a unified design space over query tokens, key aggregation, and cross-sample aggregation. We then show that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads. Across eleven VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect. Beyond replicating the causal-sparse-universal triad established for text retrieval heads, VRHs exhibit several properties not previously reported: they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction; they are functionally specific, preserving output format while corrupting localization; and they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.