NTH
AI research

PhyGround: Benchmarking Physical Reasoning in Generative World Models

AuthorsJuyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y. Lu, Enfu Nan, Hokin Deng, Edmund Yeh, Sarah Ostadabbas, Yun Fu, Jennifer Dy, Pu Zhao, Yanzhi Wang

May 18, 2026 2 min read
Watch on YouTube
The one-line take

PhyGround is a new benchmark and judging system for testing whether AI-generated videos really obey physical laws, with detailed human annotations and an open physics-aware evaluator.

Key results

13
Physical laws

PhyGround uses a criteria-grounded taxonomy of 13 visually observable physical laws across solid-body mechanics, fluid dynamics, and optics.

250
Benchmark prompts

The benchmark contains 250 curated text+image-to-video prompts, each augmented with an explicit expected physical outcome.

2,000
Generated videos

Eight video generation models were evaluated on the full prompt set, producing 2,000 generated videos.

459 annotators
Annotator pool

A large controlled human study recruited 459 annotators before quality control filtering.

5,796 complete annotations
Human annotations

The human study produced 5,796 complete annotations and over 37.4K fine-grained labels.

37.4K+
Fine-grained labels

Each annotation included multiple dimension and law-level labels, yielding over 37.4K fine-grained labels in total.

What the paper found

PhyGround reframes physical-reasoning evaluation for generative world models by replacing coarse “physics” scores with a criteria-grounded taxonomy of 13 visually observable laws spanning solid-body mechanics, fluid dynamics, and optics. The benchmark contains 250 text+image-to-video prompts, each augmented with an explicit expected physical outcome, and 2,000 generated videos from eight models, including Veo-3.1, Wan2.2-27B-A14B, OmniWeaving, Cosmos-Predict2.5-14B/2B, and LTX-2.3-22B/2-19B. Its main novelty is diagnostic granularity: laws such as gravity, momentum, impenetrability, continuity, reflection, and shadow are scored separately through 2–3 observable sub-questions on a 1–5 Likert scale, exposing failures that holistic metrics hide. A large controlled human study with 459 annotators produced 5,796 complete annotations and 37.4K fine-grained labels; after quality control, 352 annotators remained and split-half model-ranking reliability exceeded Spearman’s ρ = 0.90. The human results show no model exceeds 3.3/5 overall, with especially weak solid-body reasoning; for example, Wan2.2-27B-A14B and Veo-3.1 tie at 3.28 overall but differ sharply, with Wan2.2 stronger on solid-body stability while Veo-3.1 leads on fluid and optics by 0.47 and 0.14 points, respectively. To enable reproducible automation, the authors release PhyJudge-9B, an open physics-specialized VLM judge fine-tuned with LoRA on PhyGround labels; it reduces aggregate relative bias to 3.3% versus 16.6% for Gemini-3.1-Pro, and far outperforms base Qwen judges whose bias remains above 29%.

Original abstract

Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/.

Read the original paper