Visual Grounding in Zero-Shot Vision-Language Control
AuthorsJ. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà
Resources
The study shows that many vision-language controllers succeed without truly seeing, while selective hazard assistance plus deterministic perception offers a safer path to grounded control.
Key results
Total model-control calls evaluated across two embodiments and three simulation environments.
Image-only renderer-calibrated controller’s mean absolute error in meters.
Two-model symmetry-consensus guardian performance on the 272-frame holdout.
Committed balanced accuracy after abstaining on tied votes.
Coverage associated with the guardian’s 0.973 committed balanced accuracy.
Offline agreement with the renderer policy when deterministic perception retains lateral authority.
What the paper found
This paper tests whether vision-language models genuinely use visual input when controlling simulated vehicles and drones, rather than exploiting conservative actions or simulator dynamics. Across 32874 scored calls, the authors apply blind-image, repeated-input, shuffled-frame, noise, non-visual, and lane-axis reflection tests to models including Qwen2.5-VL-72B, Gemma4-12B, Qwen3.5-9B, Mistral’s Ministral 3, SmolVLM, and systems served through an OpenAI-compatible endpoint. Direct control is largely ungrounded: constant-SLOW can outperform a scripted geometric controller, several models are nearly constant or respond only to image presence, and even models that detect longitudinal hazards fail to swap LEFT and RIGHT under reflection. In contrast, an image-only renderer-calibrated controller estimates lead distance with 0.090 m mean absolute error and exact mirror equivariance, proving the interface contains sufficient visual information. A leakage-controlled symmetry-consensus guardian combining Gemma4-12B and Qwen3.5-9B reaches 0.954 balanced accuracy on a 272-frame holdout; abstaining on tied votes raises committed balanced accuracy to 0.973 at 0.824 coverage. A modular design that leaves lateral authority to deterministic perception achieves 0.934 action agreement, while model-predictive control cannot recover geometry missing from collapsed VLM intent. The central conclusion is that current VLMs are better treated as selective longitudinal hazard assistants than as monolithic zero-shot controllers, and that grounding evaluations must separate image use, spatial understanding, and control validity.
Original abstract
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.