NTH

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

AuthorsChang-Bin Zhang, Yujie Zhong, Qiang Zhang, Kai Han

June 26, 2026 2 min read
Watch on YouTube
The one-line take

iVGR teaches multimodal language models to keep visual grounding inside their reasoning instead of forcing explicit boxes at inference time, improving fine-grained visual reasoning with reinforcement learning.

Key results

76.6
Qwen2.5-VL-7B avg

Average accuracy achieved by iVGR-Qwen2.5-VL-7B

80.6
Qwen3-VL-8B avg

Average accuracy achieved by iVGR-Qwen3-VL-8B

82.9
Qwen3-VL-32B avg

Average accuracy achieved by iVGR-Qwen3-VL-32B

What the paper found

iVGR, from Kai Han’s group at the University of Hong Kong with coauthors Chang-Bin Zhang, Yujie Zhong, and Qiang Zhang, is a reinforcement-learning framework for multimodal large language models that internalizes visually grounded reasoning instead of forcing explicit bounding boxes at inference. The paper shows that for off-the-shelf grounded models such as DeepEyes-7B and TreeVGR-7B, explicit grounded CoT can be worse than standard textual CoT, motivating a dual-stream training design: one rollout stream learns grounded reasoning with box rewards, while a second textual stream is aligned to the best grounded trajectories through a novel consistency reward scored by an external judge model, Qwen2.5-72B-Instruct. Trained on a 35K cold-start mix and a 37K RL dataset expanded with 14K general reasoning samples, iVGR improves Qwen2.5-VL-7B from 70.1 to 76.6 average accuracy, lifting V* from 78.5 to 86.4 and HR8K from 65.1 to 75.5; on Qwen3-VL-8B it reaches 80.6 average, and on Qwen3-VL-32B it reaches 82.9. The method also scales at test time: adding predicted crops and a union crop raises Qwen3-VL-8B to 85.6 average on high-resolution benchmarks, with HR4K at 84.3 and V* at 93.2. Ablations show the consistency reward is the key driver, adding 7.3 points over the baseline, while the rollout archive further stabilizes training and raises the average to 72.4 in the controlled study.

Original abstract

While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains underexplored. In this work, we empirically find that mandating explicit object boxes in visually grounded CoT during inference often degrades performance compared to standard textual CoT, which reasons without explicit visual grounding. We hypothesize that the visual localization capability can be internalized into the textual CoT and that the mandatory explicit grounding introduces unnecessary interference with the model's primary objective of answer prediction. To address this problem, we propose Internalizing Visually Grounded Reasoning (\textbf{iVGR}), a novel reinforcement learning framework that transfers localization capabilities into the textual reasoning process. We employ a dual-stream training strategy, where a textual stream is aligned with a high-quality visually grounded stream via a proposed consistency reward, enabling the model to localize accurately without explicit grounding during inference. Extensive experiments demonstrate that our method significantly outperforms existing baselines on fine-grained benchmarks, while maintaining the flexibility to support tool-assisted inference workflows.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis