NTH

From Generation to Simulation: How Far Are World Models from Being True Simulators?

AuthorsTong Wang, Huan Deng, Mucheng Yang, Yang He, Xiaohui Kuang, Gang Zhao

August 29, 2026 2 min read
Watch on YouTube
The one-line take

This study asks whether today’s generative world models can truly replace simulators and finds that they still lack reliable physics, state feedback, and stable long-horizon behavior.

Key results

200
Audited corpus

Representative works evaluated across eight simulator capabilities

6
Runtime state-feedback interfaces

Implementation papers exposing queryable entity states or physical parameters, out of 163

45
Sensor-level feedback

Implementation papers exposing outputs beyond RGB, out of 163

62.5%
Controllability coverage

Share of the corpus treating controllability as a principal contribution

15
V-JEPA 2 planning advantage

Times faster than comparable pixel-generation planning

What the paper found

This capability audit asks whether modern generative world models are true simulators or merely convincing generators. It evaluates 200 representative works against eight simulator requirements: asset construction, physics, interaction, controllability, stability, state feedback, diversity, and evaluation. The field’s strongest areas are controllability and interaction, illustrated by systems such as DeepMind’s Genie, NVIDIA’s Cosmos platform, and interactive video models related to Sora; however, these systems generally sample plausible visual futures rather than enforce invariant physical laws. The most serious deficiency is state feedback: across 163 implementation papers, only 6 expose a runtime interface for queryable entity states or physical parameters, while 45 provide sensor-level outputs beyond RGB. This limits direct use by controllers, reinforcement-learning agents, and safety-critical evaluators. The audit also finds that controllability receives 62.5% coverage, whereas physics-engine functionality receives only 17%, revealing a major research imbalance. Joint-embedding models such as V-JEPA 2 offer an alternative by planning in latent space: they require 16 seconds per action compared with about 4 minutes for comparable pixel-generation planning, a 15 times speed advantage, but their representations are difficult to interpret. Techniques including autoregressive-diffusion distillation, latent-action learning, differentiable physics, explicit 3D memory, and causal-intervention benchmarks are narrowing the gap. The conclusion is conditional: world models can replace simulators in restricted scenarios, but formal physics guarantees, persistent long-horizon state, structured feedback, reproducibility, and downstream-utility evaluation remain unresolved. The proposed roadmap prioritizes formalized physics, unified action interfaces, first-class state feedback, stability, utility-based evaluation, and hybrid architectures.

Original abstract

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the remaining distance from generation to simulation lacks a systematic assessment. We present a capability-based study using an external yardstick: eight capabilities of a traditional simulator, namely asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. We trace three main technical routes--latent dynamics, video generation, and joint-embedding prediction--and map exactly 200 representative works published from 2018 to June 2026 onto these capabilities. Our analysis shows that world models have achieved functional substitution in interaction and controllability for specific scenarios, but remain short of traditional simulators in formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution. State feedback is the most neglected cross-route shortcoming: only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. We identify six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization. Project page: https://github.com/AtongWang/world-model-simulators

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis