See2Think: Do Multimodal Models Really Use Intermediate Visual States?
AuthorsSiyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
Resources
See2Think tests whether multimodal models truly reason with their sketches and visual feedback, finding that they often choose the right visual actions but struggle to render and use the results faithfully.
Key results
Visually dependent reasoning problems in the benchmark.
Overall accuracy with text-only chain-of-thought.
Best overall result for Gemini under action planning without rendering.
Drop for trajectories with Feedback Uptake score 1.
Average process score across evaluated models and environments.
What the paper found
See2Think asks whether multimodal models genuinely use the sketches, highlights, crops, and annotations they generate during reasoning, rather than merely producing plausible visual traces. The authors from Shanghai AI Laboratory, Central South University, and partner universities introduce See2ThinkBench, a benchmark of 1200 visually dependent problems spanning 12 categories across 2D structured reasoning, 3D scenes, and real-world environments, and Visual Action-of-Thought, or VAoT, which records textual thoughts, structured visual actions, externally rendered states, and subsequent reasoning. They compare four systems—OpenAI’s GPT-5.5 and GPT-o3, Google’s Gemini 3.5 Flash, and Alibaba’s Qwen3-VL-32B-Instruct—under text-only CoT, action planning without rendering, closed-loop VAoT, and corrupted-feedback VAoT. No strategy consistently wins: GPT-5.5 reaches 50.2% with CoT but 44.6% with VAoT, while Gemini 3.5 Flash performs best with action planning without rendering at 58.2%. Process diagnosis finds that models usually select relevant operations, with overall Action Relevance at 0.985, but rendering is substantially weaker, with Render Faithfulness at 0.616. Crucially, visual-state utility differs from behavioral dependence: when task-relevant feedback is corrupted, accuracy in 3D scenes falls by 15.5 percentage points for trajectories with maximal Feedback Uptake, and corrupted states change semantic answers in 32.7% to 54.2% of cases across models. The central conclusion is that “thinking with images” is conditional and fragile; faithful rendering and correct downstream use, not action selection alone, determine whether intermediate visual states help.
Original abstract
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.