Wonder: Video World Model Done Better
AuthorsJiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
Resources
Wonder turns video generation into an interactive, navigable world that remembers what it has seen and responds to camera motion in real time.
Key results
Reported real-time rate for minute-scale world exploration.
Overall visual-quality average on the image-to-video benchmark.
Translational camera-following error in image-to-video evaluation.
Rotational camera-following error in image-to-video evaluation.
Overall visual-quality average on the video-to-video benchmark.
Number of diverse images used in the image-to-video benchmark.
What the paper found
Adobe Research and Johns Hopkins University present Wonder, a real-time video world model that turns a single image into an explorable environment or re-shoots a source video from user-controlled viewpoints. Built on Wan2.1-I2V-14B, the system co-designs three components: a pixel-space coordinate field that renders camera translation, rotation, and parallax as dense visual cues; sparse full-fidelity memory that retrieves relevant historical key-value chunks with constant active attention; and a self-forcing distillation pipeline using sparse context forcing, a timestep-wise mixture of students, and low-frequency GAN control regularization. Its 4-step sampler uses three 14B student generators sharing a KV cache, while FlashAttention-3, CUDA graph-style compilation, and multi-GPU execution support minute-scale generation at 16 FPS. On an image-to-video benchmark, Wonder reaches an average VBench score of 0.8558 and records translational and rotational relative pose errors of 0.0132 and 0.0784, respectively. For video-to-video generation, it achieves an average score of 0.8527 while preserving dynamic scenes under novel camera trajectories. The model is evaluated on 1,000 images and 500 videos, and its sparse memory enables coherent revisits and stable latency as exploration history grows.
Original abstract
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.