Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
AuthorsJiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
Resources
Zero-WAM teaches robots unseen manipulation tasks by showing them human videos, bringing in-context learning from language models into embodied robotics.
Key results
Automatically generated human-robot in-context learning pairs.
Average success rate across seven unseen RoboTwin 2.0 tasks.
Absolute percentage-point improvement in average success rate.
Average success without in-context future chunk prediction, versus 46.95% with it.
What the paper found
Zero-WAM reframes zero-shot robot manipulation as in-context world-action modeling: instead of relying only on text, a robot watches a human demonstration video that visually specifies object states, action order, and long-horizon task evolution. Its causal video-action policy, initialized from Wan-2.2-TI2V-5B, predicts future robot video chunks and decodes executable actions through a Mixture-of-Transformers architecture. To scale paired supervision, the HumanGen pipeline converts task-sampled robot trajectories into semantically matched human videos using VLMs such as Google DeepMind’s Gemini 3.1 Pro and Qwen3.6-Plus, image models including Nano Banana 2 and Qwen-Image-2.0, and video generators such as Wan 2.7 and Kling AI 3.0. HumanGen contains 74.2K human-robot ICL pairs across 8.6K tasks, while task-balanced pretraining samples more than 6,000 tasks and approximately 400K robot trajectories per epoch from AgiBot, InternData-A1, Open-X-Embodiment, RoboCOIN, and RoboMIND. The key training innovation, in-context future chunk prediction, forces the policy to encode task information from the human video rather than shortcutting from recent robot history or language. On seven unseen RoboTwin 2.0 tasks, Zero-WAM reaches 46.95% average success, a 29.50 percentage-point gain over LingBot-VA. Removing the future-chunk objective reduces performance from 46.95% to 28.55%. Real-robot tests show transfer to unseen multi-object, long-horizon, and precision-insertion configurations without task-specific robot data or parameter updates.
Original abstract
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.