iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
AuthorsChang-Bin Zhang, Yujie Zhong, Qiang Zhang, Kai Han
Resources
iVGR teaches multimodal language models to keep visual grounding inside their reasoning instead of forcing explicit boxes at inference time, improving fine-grained visual reasoning with reinforcement learning.
Key results
Average accuracy achieved by iVGR-Qwen2.5-VL-7B
Average accuracy achieved by iVGR-Qwen3-VL-8B
Average accuracy achieved by iVGR-Qwen3-VL-32B
What the paper found
iVGR, from Kai Han’s group at the University of Hong Kong with coauthors Chang-Bin Zhang, Yujie Zhong, and Qiang Zhang, is a reinforcement-learning framework for multimodal large language models that internalizes visually grounded reasoning instead of forcing explicit bounding boxes at inference. The paper shows that for off-the-shelf grounded models such as DeepEyes-7B and TreeVGR-7B, explicit grounded CoT can be worse than standard textual CoT, motivating a dual-stream training design: one rollout stream learns grounded reasoning with box rewards, while a second textual stream is aligned to the best grounded trajectories through a novel consistency reward scored by an external judge model, Qwen2.5-72B-Instruct. Trained on a 35K cold-start mix and a 37K RL dataset expanded with 14K general reasoning samples, iVGR improves Qwen2.5-VL-7B from 70.1 to 76.6 average accuracy, lifting V* from 78.5 to 86.4 and HR8K from 65.1 to 75.5; on Qwen3-VL-8B it reaches 80.6 average, and on Qwen3-VL-32B it reaches 82.9. The method also scales at test time: adding predicted crops and a union crop raises Qwen3-VL-8B to 85.6 average on high-resolution benchmarks, with HR4K at 84.3 and V* at 93.2. Ablations show the consistency reward is the key driver, adding 7.3 points over the baseline, while the rollout archive further stabilizes training and raises the average to 72.4 in the controlled study.
Original abstract
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains underexplored. In this work, we empirically find that mandating explicit object boxes in visually grounded CoT during inference often degrades performance compared to standard textual CoT, which reasons without explicit visual grounding. We hypothesize that the visual localization capability can be internalized into the textual CoT and that the mandatory explicit grounding introduces unnecessary interference with the model's primary objective of answer prediction. To address this problem, we propose Internalizing Visually Grounded Reasoning (\textbf{iVGR}), a novel reinforcement learning framework that transfers localization capabilities into the textual reasoning process. We employ a dual-stream training strategy, where a textual stream is aligned with a high-quality visually grounded stream via a proposed consistency reward, enabling the model to localize accurately without explicit grounding during inference. Extensive experiments demonstrate that our method significantly outperforms existing baselines on fine-grained benchmarks, while maintaining the flexibility to support tool-assisted inference workflows.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.