NTH

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

AuthorsYuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen

August 29, 2026 2 min read
Watch on YouTube
The one-line take

PAWBench asks whether video generators merely make plausible futures or actually reproduce the full range and probabilities of how the world can unfold.

Key results

50
PAWBench scenarios

Diagnostic scenarios spanning eight physical mechanism groups.

11
Evaluated video models

Current video generators benchmarked across PAW-Calibration and PAW-Coverage.

20.5
Cosmos 3 Super I2V Calibration TVD

Best average PAW-Calibration TVD, reported on the ×100 scale.

71.7%
LTX-2.3 Coverage

Highest baseline average valid-support recovery.

74.8%
LTX-2.3 with C2C Coverage

Coverage after Couple to Control coupled-noise sampling.

What the paper found

PAWBench argues that a video world model must do more than generate one plausible continuation: under the same observation and action, it should reproduce both the valid futures and their probabilities. The benchmark contains 50 scenarios across eight physical mechanism groups, split between PAW-Calibration, which compares repeated-rollout frequencies with known references using total variation distance, and PAW-Coverage, which measures recovery of valid outcomes. PAWEval uses Gemini 3.5 Flash to convert repeated videos into terminal-outcome distributions. Across 11 systems, including Google’s Veo3.1 Fast, ByteDance’s Seedance 2, OpenAI-related world-model comparisons, NVIDIA’s Cosmos 3 Super I2V, Wan2.2, and MiniMax H3, no model reliably achieved calibrated probabilities, broad coverage, and consistent scene-level readability. With 50 rollouts per scene, Cosmos 3 Super I2V had the best Calibration TVD at 20.5, while LTX-2.3 reached the highest Coverage average at 71.7 percent, but neither led on every requirement. The gap was not explained by sampling noise: observed average TVD was 31.2, versus a matched reference-sampling baseline whose 99th-percentile average stayed below 9.22. Coupled initial noise through Couple to Control improved exploration, raising LTX-2.3 Coverage to 74.8 percent, but did not change the learned distribution. LoRA training on different left-right outcome mixtures shifted probabilities without learning scene-conditioned physics. Overall, prompting, sampling, and fine-tuning can steer individual outcomes, but current generators remain far from probabilistically aligned world modeling.

Original abstract

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis