NTH

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

AuthorsZhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu

AffiliationsDalian University of Technology · Nanyang Technological University

September 25, 2026 2 min read
Watch on YouTube
The one-line take

VA-Bench tests whether vision-language models can actively gather the right views and turn spatial understanding into reliable robot actions.

Key results

14
VABench task families

Base benchmark families covering single-arm and dual-arm manipulation.

12
Evaluated model conditions

Primary model conditions tested across the benchmark.

53.93%
Best macro-average task success

Three-run success average, with ±3.17% run-to-run variation.

57.50%
Active-view success

Qwen3.8-max performance with active camera control.

27.86%
Passive-view success

Matched Qwen3.8-max performance with five predefined views.

61/100
Long-horizon object placements

Best partial progress achieved by Qwen3.8-max; strict episode success was 0.00%.

What the paper found

VABench evaluates whether multimodal large language models can complete the full embodied loop of observing, reasoning, acting, and revising, rather than merely answering spatial questions. In the RoboTwin-based benchmark, models interpret RGB-only task demonstrations, actively select camera viewpoints, infer robot–object geometry, issue metric Cartesian commands, and respond to execution feedback, without privileged object poses, depth, oracle trajectories, or learned action heads. The suite contains 14 task families spanning single-arm grasping, placement, tool use, and dual-arm coordination, plus held-out geometry variants and a five-object long-horizon track. Across 12 model conditions—including Qwen3.8-max, Opus-5, GPT-5.6-sol, Gemini-3.6, and Doubao-2.1—the best three-run macro-average task success is 53.93 ± 3.17%, despite target-localization diagnostics reaching 100.0%. Active perception is central: for Qwen3.8-max, replacing viewpoint selection with five passive views reduces success from 57.50% to 27.86%. Transfer to changed geometry can cost 32.14 percentage points, and strict long-horizon success remains 0.00%, even though Qwen3.8-max completes 61/100 object placements. The benchmark’s nine trajectory diagnostics expose failures in spatial relations, fine-grained pre-contact analysis, and online correction. Its harness supports OpenAI-compatible and Anthropic endpoints, but keeps execution model-agnostic, isolating the spatial grounding and control abilities of systems such as GPT and Gemini.

Original abstract

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis