Flash-WAM: Modality-Aware Distillation for World Action Models
AuthorsArman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang
Resources
Flash-WAM makes diffusion-based robot world models fast enough for real-time control by using different distillation parameterizations for video and action streams, cutting inference from seconds to milliseconds without losing much task success.
Key results
Average success when the standard consistency function is applied uniformly to both video and action streams.
LingBot-VA baseline average success on RoboTwin 2.0.
Per-chunk latency in ms on a single NVIDIA L40S after Flash-WAM distillation.
Latency reduction over the LingBot-VA teacher on RoboTwin 2.0.
Average success at 1v/2a on RoboTwin 2.0.
Average success at 1v/2a across Spatial, Object, Goal, and Long-horizon suites.
What the paper found
Flash-WAM introduces a modality-aware distillation strategy for world action models, targeting the core failure mode of naive consistency distillation in joint video-action diffusion. The paper shows that LingBot-VA’s asymmetric noise schedules place video in a high-noise regime and actions in a low-noise regime, so a single consistency function suppresses action gradients and can collapse RoboTwin 2.0 success from 91.25% to 23.97%. Flash-WAM fixes this by pairing a Karras-style variance-preserving consistency function for video with a linear-gradient-scaling parametrization for actions, then training both heads jointly with a single shared transformer. Instantiated on LingBot-VA, the method compresses inference from 25 video steps and 50 action steps to 1 step per modality, cutting per-chunk latency on an NVIDIA L40S from 8.1 seconds to 348 ms, a 23.3× speedup. It preserves 85.54% average success on RoboTwin 2.0 and 95.7% on LIBERO at 1v/2a, and still reaches 81.41% and 95.1% at the more aggressive 1v/1a budget. On a Unitree G1 humanoid robot, Flash-WAM achieves 60.0% average success across three manipulation tasks, clearly outperforming Video-only LCM and reduced-NFE inference without distillation. The main technical contribution is not a new architecture, but a structural analysis showing that modality-specific noise regimes require modality-specific consistency functions to recover real-time control without sacrificing task success.
Original abstract
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from $8.1$ seconds to $348$ ms on NVIDIA L40S, a $23{\times}$ speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks ($85.5\%$ RoboTwin 2.0, $95.7\%$ LIBERO) and substantially recovers real-world performance ($60\%$ average on a Unitree G1 humanoid robot), while naive consistency distillation drops to $24\%$ at the same step budget.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.