NTH

Mental World Modeling

AuthorsHao Fei, Yiran Zhao

August 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that truly useful world models must simulate not only what happens in a scene, but also what the people in it believe, want, and intend.

Key results

448
Menti-Bench records

Process-annotated situated decision records spanning text, image, and sounding-video stories.

87.9
M ENTIS average F1

Average final-action F1 across eight OpenAI and Anthropic world models.

12.1
Mental-channel ablation loss

F1-point decrease when mental state and mental observation are removed.

16.5
Physical-channel ablation loss

F1-point decrease when physical state and physical observation are removed.

3.5
Gold-transition gain

F1-point improvement from replacing predicted successor states with gold transitions.

98.5
Human reference F1

Human performance under the identical evaluation protocol.

What the paper found

In “Mental World Modeling,” Hao Fei of the University of Oxford and Yiran Zhao of the National University of Singapore argue that conventional world models predict physical change but miss the beliefs, goals, intentions, emotions, and social norms that drive human action. Their framework, Mental World Modeling, represents a coupled physical-mental state, renders a target-specific partial observation, decomposes actions into physical carriers and mental meanings, and simulates joint successor states. The authors implement this design in M ENTIS, a training-free, inspectable pipeline for parsing scenes, generating observations, simulating candidate branches, and evaluating physical plausibility, mental consistency, and social appropriateness. On Menti-Bench, a 448-record testbed covering text, images, and sounding video, M ENTIS averages 87.9 F1 across eight world models: five OpenAI models, including GPT-5.6-sol and GPT-4.1, and three Anthropic models, including Claude Opus 4-8. Removing mental information costs 12.1 F1 points, removing physical information costs 16.5, and replacing coupled transitions with independent ones costs 6.4, showing that both channels and their interaction matter. The largest oracle improvement comes from transition simulation, which adds 3.5 F1 points. Gains are strongest in interpersonal scenes, where M ENTIS improves over direct answering by 26.4 F1 points. Humans reach 98.5 F1, leaving transition modeling and state parsing as the main research bottlenecks.

Original abstract

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis