OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
AuthorsYuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao
Resources
OpenWAM turns world-action model training into an open, modular science experiment and shows how video knowledge and robot experience can combine to improve control across simulations and real robots.
Key results
Frames used to pretrain OpenWAM-α.
Total hours of egocentric, real-robot, and synthetic-robot data.
Shared action representation across robot embodiments.
Average success rate under clean-to-randomized transfer.
Overall success rate on the mobile manipulation benchmark.
Average success across six real-world single-arm tasks.
What the paper found
OpenWAM reframes World–Action Model research as a controlled, modular engineering problem rather than a collection of monolithic systems. OpenWAM-Infra separates visual encoders, video and action backbones, visibility masks, training, deployment, and evaluation, allowing systematic comparisons across six architecture variants. The controlled OpenWAM-Study finds that effective transfer from video-generation priors requires a capable backbone such as Wan2.2-TI2V-5B and a compact, information-rich latent space; world–action synergy requires dedicated action capacity, explicit world-to-action attention, and synchronized joint denoising. Embodied pretraining mainly improves out-of-distribution generalization, with egocentric video supplying visual diversity and robot trajectories supplying executable action grounding. The resulting OpenWAM-α uses a dual-system architecture with a Wan2.2-TI2V-5B video stream, a 1B-parameter ActionDiT, mutual attention, and an 80-D unified action space. It is pretrained jointly on 518.5M frames totaling 6,369 hours from human egocentric, real-robot, and synthetic-robot data. OpenWAM-α scores 69.0 average success on RoboTwin2.0-Clean2Random and 49.4 success rate on EBench, while reaching 82.5% average success across six single-arm real-robot tasks. Compared with vision-language-action systems such as Qwen-RobotManip and NVIDIA’s GR00T, the results suggest WAMs fit in-distribution tasks strongly but remain more vulnerable to visual distribution shifts, especially when pixel-level future prediction accumulates error. The complete open stack, pretrained model, evaluation protocols, and data recipes are released for reproducible research.
Original abstract
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.