NTH

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

AuthorsJisong Cai, Long Ling, Shiwei Chu, Zhongshan Liu, Jiayue Kang, Zhixuan Liang, Wenjie Xu, Yinan Mao, Weinan Zhang, Xiaokang Yang, Ru Ying, Ran Zheng, Yao Mu

June 19, 2026 2 min read
Watch on YouTube
The one-line take

AHA-WAM speeds up robot control by letting one model plan the long-horizon world at a slower pace while another executes actions quickly using routed context from the planner.

Key results

92.80%
RoboTwin 2.0 average success

AHA-WAM mean success across 50 RoboTwin 2.0 tasks

78.3%
Real-world task success

Average success across 4 physical manipulation tasks

24.17
Closed-loop frequency

AHA-WAM action update rate in Hz

41.37
Latency

AHA-WAM end-to-end action-chunk latency in ms

56.95
AHA-WAM-Flash frequency

Distilled fast sampler control rate in Hz

17.56
AHA-WAM-Flash latency

Distilled action-chunk latency in ms

What the paper found

AHA-WAM, from Shanghai Jiao Tong University, Shanghai AI Laboratory, Baidu AI Cloud, and The University of Hong Kong, reframes world-action modeling for robot manipulation as an asynchronous two-timescale system: a low-frequency video DiT acts as a long-horizon world planner, while a high-frequency action DiT runs closed-loop control on short action chunks. Its key novelty is Observation-Guided Video-Context Routing, which uses the latest visual observation to adapt cached planner context without rerunning the video model, plus horizon-adaptive offset training that makes the action policy robust to planner-executor phase misalignment. Built on Wan2.2-5B and trained without robot-data pretraining, the system reaches 92.80% average success on RoboTwin 2.0 across 50 tasks and 78.3% success on four real-world bimanual tasks, outperforming Fast-WAM and matching or exceeding strong VLA baselines in deployment. The asynchronous schedule also raises closed-loop control from 5.26 Hz to 24.17 Hz with 41.37 ms latency, and the distilled AHA-WAM-Flash variant pushes this to 56.95 Hz at 17.56 ms, showing that long-horizon visual dynamics can be retained while sharply reducing control latency.

Original abstract

World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, while a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to real-time execution state without rerunning the video DiT. Experiments on RoboTwin and real-world manipulation tasks show that AHA-WAM achieves state-of-the-art performance without any robot-data pretraining, attaining 92.80% average success on RoboTwin and 78.3% success across 4 real-world tasks, while reaching 24.17 Hz closed-loop control with a 4.59x speedup over Fast-WAM.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis