NTH

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

AuthorsYuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang

August 12, 2026 2 min read
Watch on YouTube
The one-line take

WorldExam tests whether video models merely look convincing or can actually understand scenes well enough to react plausibly to what happens in them.

Key results

1,474
WorldExam cases

Benchmark size across eight evaluation tasks.

20
Evaluated models

Representative camera-, action-, and language-driven models.

97.33
NeoVerse Camera Control

Highest reported camera-control score in the static-scene results.

85.10
Veo 3.1 Social Interaction

Leading score for plausible responses from nearby agents.

0.8614
Checklist-human Spearman correlation

Overall agreement between GPT-5.5 checklist judgments and human evaluation.

What the paper found

WorldExam evaluates video world models not only by visual quality and instruction following, but by inherent reactivity: whether a generated world infers and depicts consequences that the input leaves unstated. Its hierarchical benchmark contains 1,474 cases across eight tasks and four levels—Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity—using interface adaptation for camera-driven SE(3) trajectories, action-driven discrete controls, and language prompts. The benchmark evaluates 20 models in static-scene and dynamic-interaction tracks, combining VGGT-Ω and SAM2 for 3D trajectory analysis with GPT-5.5 checklist-based judgments of terrain, object, social, physical, and goal-directed reactions. Results expose a capability split: NeoVerse reaches a 97.33 Camera Control score, while action-driven WorldPlay reaches 92.74; language-driven systems are more reactive, with Veo 3.1 scoring 85.10 on Social Interaction, 75.96 on Object Interaction, and 85.30 on Goal Completion. However, no model combines broad coverage with consistently strong performance. The study’s central finding is that visual fidelity and explicit control adherence do not imply causal scene responsiveness: models may move a subject correctly while leaving stairs, obstacles, nearby agents, or physical processes unchanged. GPT-5.5 checklist scores show strong overall human agreement, with Spearman correlation 0.8614, supporting the benchmark’s diagnostic use.

Original abstract

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis