NTH

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

AuthorsJiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu

August 28, 2026 2 min read
Watch on YouTube
The one-line take

Zero-WAM teaches robots unseen manipulation tasks by showing them human videos, bringing in-context learning from language models into embodied robotics.

Key results

74.2K
HumanGen pairs

Automatically generated human-robot in-context learning pairs.

46.95%
RoboTwin unseen-task success

Average success rate across seven unseen RoboTwin 2.0 tasks.

29.50
Gain over LingBot-VA

Absolute percentage-point improvement in average success rate.

28.55%
IFP ablation

Average success without in-context future chunk prediction, versus 46.95% with it.

What the paper found

Zero-WAM reframes zero-shot robot manipulation as in-context world-action modeling: instead of relying only on text, a robot watches a human demonstration video that visually specifies object states, action order, and long-horizon task evolution. Its causal video-action policy, initialized from Wan-2.2-TI2V-5B, predicts future robot video chunks and decodes executable actions through a Mixture-of-Transformers architecture. To scale paired supervision, the HumanGen pipeline converts task-sampled robot trajectories into semantically matched human videos using VLMs such as Google DeepMind’s Gemini 3.1 Pro and Qwen3.6-Plus, image models including Nano Banana 2 and Qwen-Image-2.0, and video generators such as Wan 2.7 and Kling AI 3.0. HumanGen contains 74.2K human-robot ICL pairs across 8.6K tasks, while task-balanced pretraining samples more than 6,000 tasks and approximately 400K robot trajectories per epoch from AgiBot, InternData-A1, Open-X-Embodiment, RoboCOIN, and RoboMIND. The key training innovation, in-context future chunk prediction, forces the policy to encode task information from the human video rather than shortcutting from recent robot history or language. On seven unseen RoboTwin 2.0 tasks, Zero-WAM reaches 46.95% average success, a 29.50 percentage-point gain over LingBot-VA. Removing the future-chunk objective reduces performance from 46.95% to 28.55%. Real-robot tests show transfer to unseen multi-object, long-horizon, and precision-insertion configurations without task-specific robot data or parameter updates.

Original abstract

Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis