WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
AuthorsKeming Wu, Yijing Cui, Wenhan Xue, Qijie Wang, Xuan Luo, Zhiyuan Feng, Zuhao Yang, Sudong Wang, Sicong Jiang, Haowei Zhu, Zihan Wang, Ping Nie, Wenhu Chen, Bin Wang
This paper introduces a benchmark to test whether video generators can actually reason about how the world should change over time, not just produce realistic-looking footage.
Key results
The benchmark evaluates image-to-video generators on 436 curated test cases with structured QA annotations across four reasoning dimensions and 22 subcategories.
On the main benchmark, the strongest closed-source generator reaches an overall Process-aware Reasoning Verification score of 39.8.
On the same benchmark, the strongest closed-source generator reaches an overall multi-dimensional quality score of 59.4.
The process-aware metric ScorePR, using Qwen3.5-27B, matches expert Human Elo with Spearman correlation 0.955 and outperforms the pairwise VLM judge.
What the paper found
WorldReasonBench reframes image-to-video evaluation as world-state prediction: given an initial frame and an instruction, can a generator produce a future clip that is physically, socially, logically, and informationally consistent rather than merely photorealistic? The benchmark contains 436 curated cases with 5–7 structured QA pairs per case across four reasoning dimensions and 22 subcategories, and it is paired with WorldRewardBench, a preference set of about 6,000 expert-annotated pairs over 1,432 videos from 11 generators. Its key methodological novelty is a two-part, human-aligned evaluation stack: Process-aware Reasoning Verification separates static outcome accuracy from dynamic process fidelity, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics with the weighted aggregate S(v)=0.4sr+0.3sc+0.3sa. On the main benchmark, closed-source systems such as Seedance2.0, Veo3.1-Fast, Kling, Wan2.6, and Sora2 outperform open-source models by roughly 2× on both ScorePR and S(v); for example, the best closed model reaches ScorePR 39.8 and S(v) 59.4, while open-source models range around ScorePR 14.4–19.6 and S(v) 21.3–38.5. The strongest automatic judge, Qwen3.5-27B with the paper’s process-aware metric, matches expert Human Elo with Spearman ρ=0.955, exceeding pairwise VLM-judge ranking at ρ=0.804. The hardest failure modes are Logic Reasoning and Information-Based reasoning, where visually plausible videos often break causal mechanisms or exact text/data preservation, revealing a persistent gap between visual polish and genuine world modeling.
Original abstract
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
Read the original paper