PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
AuthorsYuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
Resources
PAWBench asks whether video generators merely make plausible futures or actually reproduce the full range and probabilities of how the world can unfold.
Key results
Diagnostic scenarios spanning eight physical mechanism groups.
Current video generators benchmarked across PAW-Calibration and PAW-Coverage.
Best average PAW-Calibration TVD, reported on the ×100 scale.
Highest baseline average valid-support recovery.
Coverage after Couple to Control coupled-noise sampling.
What the paper found
PAWBench argues that a video world model must do more than generate one plausible continuation: under the same observation and action, it should reproduce both the valid futures and their probabilities. The benchmark contains 50 scenarios across eight physical mechanism groups, split between PAW-Calibration, which compares repeated-rollout frequencies with known references using total variation distance, and PAW-Coverage, which measures recovery of valid outcomes. PAWEval uses Gemini 3.5 Flash to convert repeated videos into terminal-outcome distributions. Across 11 systems, including Google’s Veo3.1 Fast, ByteDance’s Seedance 2, OpenAI-related world-model comparisons, NVIDIA’s Cosmos 3 Super I2V, Wan2.2, and MiniMax H3, no model reliably achieved calibrated probabilities, broad coverage, and consistent scene-level readability. With 50 rollouts per scene, Cosmos 3 Super I2V had the best Calibration TVD at 20.5, while LTX-2.3 reached the highest Coverage average at 71.7 percent, but neither led on every requirement. The gap was not explained by sampling noise: observed average TVD was 31.2, versus a matched reference-sampling baseline whose 99th-percentile average stayed below 9.22. Coupled initial noise through Couple to Control improved exploration, raising LTX-2.3 Coverage to 74.8 percent, but did not change the learned distribution. LoRA training on different left-right outcome mixtures shifted probabilities without learning scene-conditioned physics. Overall, prompting, sampling, and fine-tuning can steer individual outcomes, but current generators remain far from probabilistically aligned world modeling.
Original abstract
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.