Resources
This paper argues that truly useful world models must simulate not only what happens in a scene, but also what the people in it believe, want, and intend.
Key results
Process-annotated situated decision records spanning text, image, and sounding-video stories.
Average final-action F1 across eight OpenAI and Anthropic world models.
F1-point decrease when mental state and mental observation are removed.
F1-point decrease when physical state and physical observation are removed.
F1-point improvement from replacing predicted successor states with gold transitions.
Human performance under the identical evaluation protocol.
What the paper found
In “Mental World Modeling,” Hao Fei of the University of Oxford and Yiran Zhao of the National University of Singapore argue that conventional world models predict physical change but miss the beliefs, goals, intentions, emotions, and social norms that drive human action. Their framework, Mental World Modeling, represents a coupled physical-mental state, renders a target-specific partial observation, decomposes actions into physical carriers and mental meanings, and simulates joint successor states. The authors implement this design in M ENTIS, a training-free, inspectable pipeline for parsing scenes, generating observations, simulating candidate branches, and evaluating physical plausibility, mental consistency, and social appropriateness. On Menti-Bench, a 448-record testbed covering text, images, and sounding video, M ENTIS averages 87.9 F1 across eight world models: five OpenAI models, including GPT-5.6-sol and GPT-4.1, and three Anthropic models, including Claude Opus 4-8. Removing mental information costs 12.1 F1 points, removing physical information costs 16.5, and replacing coupled transitions with independent ones costs 6.4, showing that both channels and their interaction matter. The largest oracle improvement comes from transition simulation, which adds 3.5 F1 points. Gains are strongest in interpersonal scenes, where M ENTIS improves over direct answering by 26.4 F1 points. Humans reach 98.5 F1, leaving transition modeling and state parsing as the main research bottlenecks.
Original abstract
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.