WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
AuthorsPeterson Co, Sicheng Hu, Chunxuan Jiao, Hongyang Cheng, Yulin Luo, Yijie Xu, Sixiang Chen, Zhongxia Zhao, Zihao Wang, DaFeng Chi, Peidong Liu, YuTong Chen, Henghua Liu, Zhihao Yuan, Huizhu Jia, Yuzheng Zhuang, Tianle Zhang, Liang Lin, Huajie Tan, Shanghang Zhang
Resources
WorldSimProbe checks whether embodied world models truly respond to actions like physical simulators, rather than merely generating plausible-looking videos.
Key results
Controlled evaluation instances across RoboTwin, ManiSkill, and LIBERO.
Mean pairwise Spearman correlation across model rankings.
Average performance on unsupported, appearance-induced contact cases.
Downstream policy success using Ctrl-World synthetic data under out-of-distribution controls.
Downstream policy success using Cosmos-3-Nano synthetic data under out-of-distribution controls.
What the paper found
WorldSimProbe introduces the Observable Simulator Contract, requiring an action-conditioned world model to produce motion corresponding to supplied controls and environment responses grounded in that realized motion, rather than merely plausible video or task success. Its five suites test local action calibration, global trajectory coverage, source-specific behavior preservation, interaction grounding, and interaction dynamics using MSE calibration, Robot-Masked Flow Alignment, TAPNext++, RobotSeg, DPFlow, and Qwen3-VL-8B as a primitive judge. Across 18608 controlled instances from RoboTwin, ManiSkill, and LIBERO, the benchmark evaluates six open-source models—IRASim, Ctrl-World, BWM, DreamDojo, LingBot-VA, and Cosmos-3-Nano—on NVIDIA H100 hardware. Fidelity declines as counterfactual motion mismatch increases, with mean Spearman correlation ρ = -0.433, while cross-platform model rankings are moderately consistent at ρ = 0.695. Interaction grounding is particularly weak for unsupported false contact, averaging 38.6%, and interaction dynamics reveal a shared failure pattern: shake succeeds only 0.0–1.2% of the time. Downstream synthetic-data experiments show similar standard-control success but sharply different out-of-distribution performance, with Ctrl-World reaching 53% and Cosmos-3-Nano 21%, demonstrating that WorldSimProbe exposes simulator failures hidden by conventional rollout evaluation.
Original abstract
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.