HumanCLAW: Can Vision-Language Models Act Through a Body?
AuthorsSiyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
Resources
HumanCLAW tests whether vision-language models truly know how to act in the physical world and finds that they struggle mainly to track their own body and progress.
Key results
Egocentric find-navigate-interact episodes in the benchmark
Scenes used to construct HumanCLAW-Bench
Gemini-3.1 success rate for finding, navigating to, and sitting on the target
Achieved-to-commanded walking displacement ratio
Navigation success after removing the skill-specific verifier
What the paper found
HumanCLAW, developed by researchers at Meta, Nanyang Technological University, and the University of Washington, asks whether vision-language models can make reliable moment-to-moment decisions through a humanoid body. Its key contribution is a decoupled evaluation loop: a frozen, off-the-shelf VLM selects one atomic skill, such as walking or turning, a skill-conditioned motion generator converts it into continuous full-body motion, and a half-physics simulator preserves gravity, collisions, contact, and object displacement while removing balance and motor-tracking failures. The resulting HumanCLAW-Bench contains 1,218 egocentric find-navigate-interact episodes across 41 indoor scenes, evaluated with nine VLMs including Gemini-3.1, GPT-5.5, Claude-4.8, and Gemma-4-31B. None solves the benchmark; Gemini-3.1 achieves the best complete-task success at 16.8%. The models generally recognize targets once they appear, but fail to estimate their own position, detect arrival, notice collisions, and place the body correctly for sitting—evidence of missing embodied self-awareness rather than inadequate object recognition. The motion layer uses a 38M-parameter, 10-layer diffusion transformer trained on AMASS motion chunks, with plug-and-play per-skill ControlNet adapters; commanded walking achieves an execution ratio of 0.966. A verifier is crucial: on a 100-episode ablation, removing it reduces navigation success from 27.0% to 2.0%, showing that compact, skill-specific checks can stabilize closed-loop action without retraining the VLM.
Original abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.