WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
AuthorsYuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
Resources
WorldExam tests whether video models merely look convincing or can actually understand scenes well enough to react plausibly to what happens in them.
Key results
Benchmark size across eight evaluation tasks.
Representative camera-, action-, and language-driven models.
Highest reported camera-control score in the static-scene results.
Leading score for plausible responses from nearby agents.
Overall agreement between GPT-5.5 checklist judgments and human evaluation.
What the paper found
WorldExam evaluates video world models not only by visual quality and instruction following, but by inherent reactivity: whether a generated world infers and depicts consequences that the input leaves unstated. Its hierarchical benchmark contains 1,474 cases across eight tasks and four levels—Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity—using interface adaptation for camera-driven SE(3) trajectories, action-driven discrete controls, and language prompts. The benchmark evaluates 20 models in static-scene and dynamic-interaction tracks, combining VGGT-Ω and SAM2 for 3D trajectory analysis with GPT-5.5 checklist-based judgments of terrain, object, social, physical, and goal-directed reactions. Results expose a capability split: NeoVerse reaches a 97.33 Camera Control score, while action-driven WorldPlay reaches 92.74; language-driven systems are more reactive, with Veo 3.1 scoring 85.10 on Social Interaction, 75.96 on Object Interaction, and 85.30 on Goal Completion. However, no model combines broad coverage with consistently strong performance. The study’s central finding is that visual fidelity and explicit control adherence do not imply causal scene responsiveness: models may move a subject correctly while leaving stairs, obstacles, nearby agents, or physical processes unchanged. GPT-5.5 checklist scores show strong overall human agreement, with Spearman correlation 0.8614, supporting the benchmark’s diagnostic use.
Original abstract
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.