NTH

Region-Level Policy Optimization for Fine-grained MLLM Perception

AuthorsYuheng Shi, Xiaohuan Pei, Minjing Dong, Chang Xu

September 19, 2026 3 min read
Watch on YouTube
The one-line take

Vision-RL2 helps multimodal language models focus on the right image regions, improving fine-grained understanding while using roughly four times fewer visual tokens.

Key results

80.1
Qwen3.5-9B six-benchmark average

Vision-RL2 score under the shared 16,384 source-image-token evaluation.

4.2
Token reduction factor

Fewer visual tokens than SD-RPN while matching its accuracy on Qwen3.5-4B.

7K
RL training pool

Question-answer pairs used to train the region proposal policy.

576
Training source-token limit

Visual-token budget used for region actions and reader rewards during training.

14.9
Gain over frozen base

Average-point improvement of the complete system in the training-aligned evaluation.

What the paper found

Region-Level Policy Optimization for Fine-grained MLLM Perception introduces Vision-RL2, a lightweight answer-aware region proposal method for multimodal large language models. Its central finding is that localization and recognition have different resolution requirements: localization remains reliable under much stronger token compression, while recognition needs dense visual evidence. Vision-RL2 starts from SD-RPN, treats coherent image regions as discrete actions, and uses a frozen MLLM reader to measure each region’s functional contribution through answer-likelihood changes when that region is removed or added. Subtractive policy optimization suppresses distracting proposals, additive optimization recovers missing evidence, and sparse visual encoding reallocates crop tokens to foreground pixels rather than background. The method updates only the proposal network, requiring no region annotations, response sampling, or reasoning trajectories. Evaluated across six fine-grained benchmarks and four backbones, including Qwen3.5, Qwen2.5-VL, and Gemma-4, Vision-RL2 consistently outperforms the base model and SD-RPN across token budgets. On Qwen3.5-9B, it reaches an 80.1 average score under a shared 16,384 source-image-token limit, and on Qwen3.5-4B it matches SD-RPN’s accuracy with 4.2-fold fewer visual tokens. Training uses 7K question-answer pairs with a 576-token source limit; under that aligned setting, the complete system improves the frozen base model by 14.9 average points. Compared with full-model approaches, including OpenAI-style image-reasoning systems and Vision-OPD, Vision-RL2 offers a more parameter-efficient routing interface while preserving fine-grained perception accuracy.

Original abstract

Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model's attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at https://github.com/YuHengsss/VisionRL2 .

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis