CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
AuthorsKechen Liu, Ola Shorinwa
Resources
CLAP aims to turn diverse human and robot videos into a shared physical simulator that can predict and control unfamiliar robot embodiments.
Key results
Continuous latent actions used to learn from unlabeled human and robot videos.
Minimum improvement over the DreamDojo-Human baseline from incorporating robot data.
Improvement over relative end-effector actions on DROID.
CLAP’s inference-time cross-policy planning success rate.
Success rate after policy finetuning inside the CLAP video world model.
What the paper found
CLAP proposes a cross-embodiment video world model that treats learned video dynamics as a zero-shot simulator across human and robotic agents, rather than restricting training to one robot morphology. Built on an SVD latent-video-diffusion backbone with CLIP language encoding, CLAP unifies control through 7-dimensional end-effector poses, templated language actions, and 32-dimensional latent actions extracted from unlabeled frame pairs. Its central curriculum, CLAP-CURR, first learns physical priors from action-free videos and then grounds them in end-effector actions for direct robot deployment. Trained on Open X-Embodiment and EgoDex, the system matches or exceeds single-embodiment performance on DROID; CLAP-CURR reaches 19.138 PSNR and 0.204 LPIPS there, while robot data improves DROID LPIPS by at least 61% over the human-video DreamDojo-Human baseline. The study also finds that absolute actions outperform relative end-effector actions, with a 14.6% LPIPS advantage on DROID, while relative actions are better for language conditioning. At inference, CLAP scores proposals from policies including π0.5 and MolmoAct-2, achieving up to 95.0% success on the red-lobster task and improving reinforcement-learning policy performance from 80.0% to 88.0% on carrot placement. Unlike general video generators such as OpenAI’s Sora or NVIDIA’s Cosmos, CLAP targets controllable physical prediction for robotics, though it still hallucinates at distribution boundaries and is not yet a replacement for calibrated physics simulators.
Original abstract
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.