NTH

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

AuthorsZhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu

August 9, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that multimodal models may appear to think with images while often ignoring the visual evidence their tools retrieve.

Key results

6
Models audited

Open visual tool-use models evaluated in the causal audit.

5
Benchmarks evaluated

Fine-grained perception benchmarks used for policy- and trajectory-level evaluation.

21.3
Mini-o3 VisualProbe policy ATE

Accuracy improvement in percentage points from tool-enabled versus direct inference.

64.2
Mini-o3 V* random-crop ATE

Accuracy-point drop under dynamic random-crop observation corruption.

20.9%
Qwen3-VL-8B Calibrated fraction

Share of V* trajectories classified as causally effective and well-planned.

7.6
Qwen3-VL-8B Calibrated contribution

Percentage-point contribution of the Calibrated subset to the V* policy-level gain.

What the paper found

Researchers at Shanghai Artificial Intelligence Laboratory and Shanghai Jiao Tong University present a causal audit of OpenAI’s “thinking-with-images” paradigm, testing whether crop-and-zoom observations genuinely influence multimodal model answers. Across 6 open models—including DeepEyes, Mini-o3, Qwen3-VL-4B/8B, Pixel Reasoner, and Thyme—and 5 fine-grained perception benchmarks, they compare tool-enabled inference with direct inference, dynamically replace every returned crop with a random crop, and perform step-level counterfactual interventions. Their key metric, Visual Evidence Gain, estimates the observation-mediated effect while cancelling action-induced shortcuts: a model may call a tool and change its answer simply because a call occurred, without using the image content. The results expose two calibration failures: “Calling Without Looking,” where observations have near-zero causal influence, and “Looking Without Planning,” where useful evidence exists but calls continue after confidence is already high or repeatedly target irrelevant regions. Aggregate gains can therefore be misleading: Mini-o3 improves VisualProbe by 21.3 percentage points, yet random-crop corruption drops its V* accuracy by 64.2 percentage points, partly through repair loops and budget exhaustion. On V*, only 20.9% of Qwen3-VL-8B trajectories are classified as Calibrated, but that minority contributes +7.6 percentage points to the overall gain. The paper argues that visual tool use should be evaluated by causal evidence use and stopping behavior, not call frequency, and hypothesizes that outcome-only reinforcement learning rewards tool-call rituals and poor credit assignment. The authors caution that these findings cover open models and crop-and-zoom, not closed systems such as OpenAI o3 or o4-mini.

Original abstract

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis