NTH

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

AuthorsRui Zhao, Kaiming Yang, Jifeng Zhu, Siyang Chen, Ziqi Wang, Weijia Wu, Kevin Qinghong Lin, Heng Wang, Mike Zheng Shou

June 11, 2026 3 min read
Watch on YouTube
The one-line take

Dream.exe tests whether video generators can dream motions that actually work in a robot simulator, turning pretty videos into a practical measure of physical understanding.

Key results

101
task suite size

Manually curated RoboCasa365 manipulation tasks in Dream.exe

3
difficulty levels

Three physical-complexity tiers in the benchmark

8
models evaluated

Video generation systems spanning closed-source, open-source, and robot-specific models

15.1%
Level 1 SR-B SeedDance 2.0

Binary task success rate at Level 1

21.4%
Level 2 SR-B Wan 2.7

Binary task success rate at Level 2

-0.03
visual-quality vs success correlation

Pearson correlation between physical plausibility and task success

What the paper found

Dream.exe, from Show Lab at the National University of Singapore, with the evaluated frontier systems including Google DeepMind’s Veo 3.1, NVIDIA’s Cosmos Policy, MiniMax’s Hailuo 2.3, Kuaishou’s Kling 3.0, ByteDance’s SeedDance 2.0, and Alibaba’s Wan 2.7, asks a direct robotics question: can the motion in a generated video be executed as a real manipulation policy? The paper builds a video-to-execution benchmark over 101 manually curated RoboCasa365 tasks spanning 3 difficulty levels, then converts each generated clip into a 7D robot trajectory using CoTracker point tracks, DVD LoRA depth estimation, Kabsch-based rotation recovery, and gripper-event inference before replaying it in MuJoCo/robosuite on a Franka Panda. Across 8 models, the key result is that visual quality is a poor proxy for physical executability: physical plausibility is essentially uncorrelated with binary task success, with Pearson r = -0.03, and Veo 3.1 can rank highest on task adherence while achieving only 3.3% Level-1 success. By contrast, general-purpose generators such as SeedDance 2.0 and Wan 2.7 reach 15.1% and 21.4% task success at Level 1 and Level 2, while the robot-specific CosmosPolicy-BenchCam peaks at 20.8% Level-1 SR-B but does not dominate task success. The strongest trajectory-following result comes from CosmosPolicy-BenchCam at 0.770 EEF-visual HSD and 0.833 DYN, yet the study shows that better geometric fidelity does not guarantee task completion. The paper’s main contribution is therefore a new grounding test for video world models: not whether a manipulation video looks realistic, but whether its implied motion can actually make a robot succeed.

Original abstract

Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follows: how well do these models reflect the physical world when their generated videos leave the screen and enter reality? We propose robotic manipulation as a concrete, measurable window onto this question: if a model has truly internalized physical laws, the motion it depicts should translate into executable robot behavior. We introduce Dream$.$exe, an evaluation framework that operationalizes this criterion through a video-to-execution pipeline. Given a scene image and a task description, Dream$.$exe synthesizes a manipulation video, converts the generated motion into robot trajectories, and executes them in a physics simulator, yielding a grounding signal that purely visual metrics cannot offer. Using this pipeline, we evaluate 8 models spanning frontier closed-source generators, open-source generators, and robot-specific models. Our benchmark covers 101 manually curated manipulation tasks at three levels of physical complexity, measured across visual quality, trajectory fidelity, and execution success. Encouragingly, several models achieve measurable execution success, suggesting that generative priors learned from internet-scale data already encode meaningful physical knowledge. Yet visual quality proves a poor predictor of executability, exposing a dimension of model capability that standard visual evaluations do not capture. Dream$.$exe will be open-sourced at https://github.com/showlab/Dream.exe.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis