AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
AuthorsJisong Cai, Long Ling, Shiwei Chu, Zhongshan Liu, Jiayue Kang, Zhixuan Liang, Wenjie Xu, Yinan Mao, Weinan Zhang, Xiaokang Yang, Ru Ying, Ran Zheng, Yao Mu
Resources
AHA-WAM speeds up robot control by letting one model plan the long-horizon world at a slower pace while another executes actions quickly using routed context from the planner.
Key results
AHA-WAM mean success across 50 RoboTwin 2.0 tasks
Average success across 4 physical manipulation tasks
AHA-WAM action update rate in Hz
AHA-WAM end-to-end action-chunk latency in ms
Distilled fast sampler control rate in Hz
Distilled action-chunk latency in ms
What the paper found
AHA-WAM, from Shanghai Jiao Tong University, Shanghai AI Laboratory, Baidu AI Cloud, and The University of Hong Kong, reframes world-action modeling for robot manipulation as an asynchronous two-timescale system: a low-frequency video DiT acts as a long-horizon world planner, while a high-frequency action DiT runs closed-loop control on short action chunks. Its key novelty is Observation-Guided Video-Context Routing, which uses the latest visual observation to adapt cached planner context without rerunning the video model, plus horizon-adaptive offset training that makes the action policy robust to planner-executor phase misalignment. Built on Wan2.2-5B and trained without robot-data pretraining, the system reaches 92.80% average success on RoboTwin 2.0 across 50 tasks and 78.3% success on four real-world bimanual tasks, outperforming Fast-WAM and matching or exceeding strong VLA baselines in deployment. The asynchronous schedule also raises closed-loop control from 5.26 Hz to 24.17 Hz with 41.37 ms latency, and the distilled AHA-WAM-Flash variant pushes this to 56.95 Hz at 17.56 ms, showing that long-horizon visual dynamics can be retained while sharply reducing control latency.
Original abstract
World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, while a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to real-time execution state without rerunning the video DiT. Experiments on RoboTwin and real-world manipulation tasks show that AHA-WAM achieves state-of-the-art performance without any robot-data pretraining, attaining 92.80% average success on RoboTwin and 78.3% success across 4 real-world tasks, while reaching 24.17 Hz closed-loop control with a 4.59x speedup over Fast-WAM.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.