Rolling-WAM: World Action Models with Rolling Imagination
AuthorsYinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
AffiliationsFudan University · Toyota Research Institute
Resources
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Key results
Rolling-WAM accelerates replanning over standard Joint-WAM inference.
Average replanning latency in milliseconds on an NVIDIA A100.
Average success rate across the four LIBERO suites.
Average success rate across clean and randomized settings.
Average success rate across three real-world humanoid manipulation tasks.
What the paper found
Rolling-WAM is a World Action Model that jointly denoises future video and robot actions, but distributes that computation across replanning cycles instead of restarting from Gaussian noise each time. It maintains a sliding window of aligned video-action chunks at staggered noise levels: the imminent chunk is fully denoised for execution, while farther chunks are progressively refined and carried forward after each new camera observation. With a default window of 5 chunks, 16 actions per chunk, and only 2 denoising steps per steady-state update, the model preserves visual imagination across an 80-action horizon. Its Mixture-of-Transformers design combines a pretrained Wan2.2-TI2V-5B video expert with an approximately 1B-parameter action transformer, using flow matching and masked cross-chunk attention. On an NVIDIA A100, Rolling-WAM reaches 215 ms replanning latency versus 978 ms for Joint-WAM, a 4.5x speedup, while retaining future-video prediction. It achieves 98.1% average success on LIBERO, 93.3% on RoboTwin 2.0, and 85.0% across three real-world Unitree G1 humanoid tasks, outperforming Joint-WAM’s 78.3% real-world average and Fast-WAM’s 75.0%. The system was evaluated alongside models such as π0.5 and GR00T N1.7, with practical relevance to low-latency physical-AI systems developed by companies including Toyota Research Institute.
Original abstract
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.
LIMBO: Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control
Jake Gonzales, Arturo Flores Alvarez, Yu-Ming Chen, Aaron D. Ames, Lillian J. Ratliff, Manikantan Nambi
LIMBO teaches humanoid robots to discover and internalize safety boundaries so they can perform agile maneuvers without relying on an online safety filter.