NTH

Rolling-WAM: World Action Models with Rolling Imagination

AuthorsYinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

AffiliationsFudan University · Toyota Research Institute

October 3, 2026 2 min read
Watch on YouTube
The one-line take

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Key results

4.5
Steady-state speedup

Rolling-WAM accelerates replanning over standard Joint-WAM inference.

215
Rolling-WAM latency

Average replanning latency in milliseconds on an NVIDIA A100.

98.1%
LIBERO success

Average success rate across the four LIBERO suites.

93.3%
RoboTwin 2.0 success

Average success rate across clean and randomized settings.

85.0%
Unitree G1 success

Average success rate across three real-world humanoid manipulation tasks.

What the paper found

Rolling-WAM is a World Action Model that jointly denoises future video and robot actions, but distributes that computation across replanning cycles instead of restarting from Gaussian noise each time. It maintains a sliding window of aligned video-action chunks at staggered noise levels: the imminent chunk is fully denoised for execution, while farther chunks are progressively refined and carried forward after each new camera observation. With a default window of 5 chunks, 16 actions per chunk, and only 2 denoising steps per steady-state update, the model preserves visual imagination across an 80-action horizon. Its Mixture-of-Transformers design combines a pretrained Wan2.2-TI2V-5B video expert with an approximately 1B-parameter action transformer, using flow matching and masked cross-chunk attention. On an NVIDIA A100, Rolling-WAM reaches 215 ms replanning latency versus 978 ms for Joint-WAM, a 4.5x speedup, while retaining future-video prediction. It achieves 98.1% average success on LIBERO, 93.3% on RoboTwin 2.0, and 85.0% across three real-world Unitree G1 humanoid tasks, outperforming Joint-WAM’s 78.3% real-world average and Fast-WAM’s 75.0%. The system was evaluated alongside models such as π0.5 and GR00T N1.7, with practical relevance to low-latency physical-AI systems developed by companies including Toyota Research Institute.

Original abstract

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis