VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
AuthorsJunxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai
Resources
VBVR-Pro is a large, verifiable playground for training and testing models that reason directly through generated images and videos.
Key results
Procedurally generated visual reasoning tasks.
Instances sampled from 250 training tasks.
Average overall improvement across nine evaluated generative models.
Fraction of repeated evaluations whose VLM-judge scores changed.
Overall VBVR-Pro-Bench score using verifiable rewards.
Throughput multiplier from the one-step-delayed verifiable-reward pipeline.
What the paper found
VBVR-Pro turns native visual reasoning—solving problems by constructing and updating images or videos—into a closed-loop research setting. Its procedurally generated suite contains 300 tasks spanning perception, spatiality, transformation, abstraction, and knowledge, with 1.25M training instances rendered as aligned video, image, and interleaved text-image trajectories. Task-specific verifiable scorers use HSV segmentation, contour detection, OCR, and trajectory tracking to check semantic constraints deterministically, avoiding the weaknesses of VLM-as-a-judge systems such as OpenAI's GPT-5.5, Google DeepMind's Gemini-3.1-Pro, and Qwen3.6-27B; repeated scoring changed in 0.0% of cases for the proposed scorer, compared with 92.8% for Gemini-3.1-Pro. Fine-tuning nine generators, including Wan2.2, ThinkMorph, and SenseNova-U1, produced an average overall gain of 0.290, with transfer improvements across RISE-Video, V-ReasonBench, RULER-Bench, MME-CoF-Pro, VideoThinkBench, BabyVision-Gen, and IntelligentVBench. The strongest trained video model, VBVR-Pro-Wan2.2-I2V-A14B, reached 66.18 on RISE-Video and 38.22 on V-ReasonBench. Controlled comparisons show video is strongest for persistent spatiotemporal state tracking, while interleaved generation is more compute-efficient; ablations indicate intermediate visual states matter more than textual chains of thought. Finally, reinforcement learning with verifiable rewards and Coefficients-Preserving Sampling reached an overall score of 0.548, outperforming VLM-reward RL at 0.508, while a one-step-delayed pipeline achieved 1.63× training throughput.
Original abstract
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.