Region-Level Policy Optimization for Fine-grained MLLM Perception
AuthorsYuheng Shi, Xiaohuan Pei, Minjing Dong, Chang Xu
Vision-RL2 helps multimodal language models focus on the right image regions, improving fine-grained understanding while using roughly four times fewer visual tokens.
Key results
Vision-RL2 score under the shared 16,384 source-image-token evaluation.
Fewer visual tokens than SD-RPN while matching its accuracy on Qwen3.5-4B.
Question-answer pairs used to train the region proposal policy.
Visual-token budget used for region actions and reader rewards during training.
Average-point improvement of the complete system in the training-aligned evaluation.
What the paper found
Region-Level Policy Optimization for Fine-grained MLLM Perception introduces Vision-RL2, a lightweight answer-aware region proposal method for multimodal large language models. Its central finding is that localization and recognition have different resolution requirements: localization remains reliable under much stronger token compression, while recognition needs dense visual evidence. Vision-RL2 starts from SD-RPN, treats coherent image regions as discrete actions, and uses a frozen MLLM reader to measure each region’s functional contribution through answer-likelihood changes when that region is removed or added. Subtractive policy optimization suppresses distracting proposals, additive optimization recovers missing evidence, and sparse visual encoding reallocates crop tokens to foreground pixels rather than background. The method updates only the proposal network, requiring no region annotations, response sampling, or reasoning trajectories. Evaluated across six fine-grained benchmarks and four backbones, including Qwen3.5, Qwen2.5-VL, and Gemma-4, Vision-RL2 consistently outperforms the base model and SD-RPN across token budgets. On Qwen3.5-9B, it reaches an 80.1 average score under a shared 16,384 source-image-token limit, and on Qwen3.5-4B it matches SD-RPN’s accuracy with 4.2-fold fewer visual tokens. Training uses 7K question-answer pairs with a 576-token source limit; under that aligned setting, the complete system improves the frozen base model by 14.9 average points. Compared with full-model approaches, including OpenAI-style image-reasoning systems and Vision-OPD, Vision-RL2 offers a more parameter-efficient routing interface while preserving fine-grained perception accuracy.
Original abstract
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model's attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at https://github.com/YuHengsss/VisionRL2 .
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.