NTH

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

AuthorsKechen Liu, Ola Shorinwa

August 29, 2026 3 min read
Watch on YouTube
The one-line take

CLAP aims to turn diverse human and robot videos into a shared physical simulator that can predict and control unfamiliar robot embodiments.

Key results

32
Latent action dimension

Continuous latent actions used to learn from unlabeled human and robot videos.

61%
DROID LPIPS improvement

Minimum improvement over the DreamDojo-Human baseline from incorporating robot data.

14.6%
Absolute-action LPIPS advantage

Improvement over relative end-effector actions on DROID.

95.0%
Best lobster-task success

CLAP’s inference-time cross-policy planning success rate.

88.0%
Carrot-task RL success

Success rate after policy finetuning inside the CLAP video world model.

What the paper found

CLAP proposes a cross-embodiment video world model that treats learned video dynamics as a zero-shot simulator across human and robotic agents, rather than restricting training to one robot morphology. Built on an SVD latent-video-diffusion backbone with CLIP language encoding, CLAP unifies control through 7-dimensional end-effector poses, templated language actions, and 32-dimensional latent actions extracted from unlabeled frame pairs. Its central curriculum, CLAP-CURR, first learns physical priors from action-free videos and then grounds them in end-effector actions for direct robot deployment. Trained on Open X-Embodiment and EgoDex, the system matches or exceeds single-embodiment performance on DROID; CLAP-CURR reaches 19.138 PSNR and 0.204 LPIPS there, while robot data improves DROID LPIPS by at least 61% over the human-video DreamDojo-Human baseline. The study also finds that absolute actions outperform relative end-effector actions, with a 14.6% LPIPS advantage on DROID, while relative actions are better for language conditioning. At inference, CLAP scores proposals from policies including π0.5 and MolmoAct-2, achieving up to 95.0% success on the red-lobster task and improving reinforcement-learning policy performance from 80.0% to 88.0% on carrot placement. Unlike general video generators such as OpenAI’s Sora or NVIDIA’s Cosmos, CLAP targets controllable physical prediction for robotics, though it still hallucinates at distribution boundaries and is not yet a replacement for calibrated physics simulators.

Original abstract

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis