NTH

HumanCLAW: Can Vision-Language Models Act Through a Body?

AuthorsSiyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo

August 4, 2026 2 min read
Watch on YouTube
The one-line take

HumanCLAW tests whether vision-language models truly know how to act in the physical world and finds that they struggle mainly to track their own body and progress.

Key results

1218
HumanCLAW-Bench episodes

Egocentric find-navigate-interact episodes in the benchmark

41
Indoor scenes

Scenes used to construct HumanCLAW-Bench

16.8%
Best complete-task success

Gemini-3.1 success rate for finding, navigating to, and sitting on the target

0.966
Walk execution ratio

Achieved-to-commanded walking displacement ratio

2.0%
Verifier ablation navigation success

Navigation success after removing the skill-specific verifier

What the paper found

HumanCLAW, developed by researchers at Meta, Nanyang Technological University, and the University of Washington, asks whether vision-language models can make reliable moment-to-moment decisions through a humanoid body. Its key contribution is a decoupled evaluation loop: a frozen, off-the-shelf VLM selects one atomic skill, such as walking or turning, a skill-conditioned motion generator converts it into continuous full-body motion, and a half-physics simulator preserves gravity, collisions, contact, and object displacement while removing balance and motor-tracking failures. The resulting HumanCLAW-Bench contains 1,218 egocentric find-navigate-interact episodes across 41 indoor scenes, evaluated with nine VLMs including Gemini-3.1, GPT-5.5, Claude-4.8, and Gemma-4-31B. None solves the benchmark; Gemini-3.1 achieves the best complete-task success at 16.8%. The models generally recognize targets once they appear, but fail to estimate their own position, detect arrival, notice collisions, and place the body correctly for sitting—evidence of missing embodied self-awareness rather than inadequate object recognition. The motion layer uses a 38M-parameter, 10-layer diffusion transformer trained on AMASS motion chunks, with plug-and-play per-skill ControlNet adapters; commanded walking achieves an execution ratio of 0.966. A verifier is crucial: on a 100-episode ablation, removing it reduces navigation success from 27.0% to 2.0%, showing that compact, skill-specific checks can stabilize closed-loop action without retraining the VLM.

Original abstract

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis