VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
AuthorsZhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu
AffiliationsDalian University of Technology · Nanyang Technological University
Resources
VA-Bench tests whether vision-language models can actively gather the right views and turn spatial understanding into reliable robot actions.
Key results
Base benchmark families covering single-arm and dual-arm manipulation.
Primary model conditions tested across the benchmark.
Three-run success average, with ±3.17% run-to-run variation.
Qwen3.8-max performance with active camera control.
Matched Qwen3.8-max performance with five predefined views.
Best partial progress achieved by Qwen3.8-max; strict episode success was 0.00%.
What the paper found
VABench evaluates whether multimodal large language models can complete the full embodied loop of observing, reasoning, acting, and revising, rather than merely answering spatial questions. In the RoboTwin-based benchmark, models interpret RGB-only task demonstrations, actively select camera viewpoints, infer robot–object geometry, issue metric Cartesian commands, and respond to execution feedback, without privileged object poses, depth, oracle trajectories, or learned action heads. The suite contains 14 task families spanning single-arm grasping, placement, tool use, and dual-arm coordination, plus held-out geometry variants and a five-object long-horizon track. Across 12 model conditions—including Qwen3.8-max, Opus-5, GPT-5.6-sol, Gemini-3.6, and Doubao-2.1—the best three-run macro-average task success is 53.93 ± 3.17%, despite target-localization diagnostics reaching 100.0%. Active perception is central: for Qwen3.8-max, replacing viewpoint selection with five passive views reduces success from 57.50% to 27.86%. Transfer to changed geometry can cost 32.14 percentage points, and strict long-horizon success remains 0.00%, even though Qwen3.8-max completes 61/100 object placements. The benchmark’s nine trajectory diagnostics expose failures in spatial relations, fine-grained pre-contact analysis, and online correction. Its harness supports OpenAI-compatible and Anthropic endpoints, but keeps execution model-agnostic, isolating the spatial grounding and control abilities of systems such as GPT and Gemini.
Original abstract
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.