NTH

Wonder: Video World Model Done Better

AuthorsJiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei

July 30, 2026 2 min read
Watch on YouTube
The one-line take

Wonder turns video generation into an interactive, navigable world that remembers what it has seen and responds to camera motion in real time.

Key results

16 FPS
Generation rate

Reported real-time rate for minute-scale world exploration.

0.8558
I2V average VBench score

Overall visual-quality average on the image-to-video benchmark.

0.0132
I2V translational RPE

Translational camera-following error in image-to-video evaluation.

0.0784
I2V rotational RPE

Rotational camera-following error in image-to-video evaluation.

0.8527
V2V average VBench score

Overall visual-quality average on the video-to-video benchmark.

1,000
I2V benchmark size

Number of diverse images used in the image-to-video benchmark.

What the paper found

Adobe Research and Johns Hopkins University present Wonder, a real-time video world model that turns a single image into an explorable environment or re-shoots a source video from user-controlled viewpoints. Built on Wan2.1-I2V-14B, the system co-designs three components: a pixel-space coordinate field that renders camera translation, rotation, and parallax as dense visual cues; sparse full-fidelity memory that retrieves relevant historical key-value chunks with constant active attention; and a self-forcing distillation pipeline using sparse context forcing, a timestep-wise mixture of students, and low-frequency GAN control regularization. Its 4-step sampler uses three 14B student generators sharing a KV cache, while FlashAttention-3, CUDA graph-style compilation, and multi-GPU execution support minute-scale generation at 16 FPS. On an image-to-video benchmark, Wonder reaches an average VBench score of 0.8558 and records translational and rotational relative pose errors of 0.0132 and 0.0784, respectively. For video-to-video generation, it achieves an average score of 0.8527 while preserving dynamic scenes under novel camera trajectories. The model is evaluated on 1,000 images and 500 videos, and its sparse memory enables coherent revisits and stable latency as exploration history grows.

Original abstract

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis