Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning
AuthorsZhikai Pan, Chih-Ting Liao, Chunrui Liu, Xi Xiao, Yitong Qiao, Chunlei Meng, Zhangquan Chen, Xin Cao
Resources
This paper tests whether LLMs can build spatial mental maps from text across languages, and finds a consistent reasoning cliff that suggests current models struggle to form robust world models.
Key results
MENTAL MAP is grounded in 100 ProcTHOR household scenes.
MENTAL MAP comprises 39 task families across the benchmark.
The benchmark is evaluated across 1,950 cells.
The human pilot reproduces the cliff with L3 equal to 0% in all eight languages.
In the human pilot, L4 is reported as around 41% with 3.3pp standard deviation across languages.
What the paper found
In “Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning,” Zhikai Pan, Chih-Ting Liao, and colleagues from the University of New South Wales introduce MENTAL MAP, a pure-text benchmark built on 100 ProcTHOR household scenes to test whether LLMs form internal spatial world models or just exploit co-occurrence. The benchmark spans 39 task families across 1,950 evaluation cells in eight typologically diverse languages—English, Mandarin, Japanese, Korean, Spanish, Arabic, Thai, and German—plus a structured-text control, and it separates static comprehension from active spatial reasoning through a six-level staircase from L0 atomic facts to L5 generative world-graph output. Across 13 models, including OpenAI’s GPT-4o, Google’s Gemini-2.5-Flash, DeepSeek-V4-Flash, and Qwen family models, the paper finds a universal “L3 cliff”: performance collapses at viewpoint transformation, with all cliff-diagnostic model-language cells scoring below half of their L0 atomic baseline and most below a quarter. The study also shows that chain-of-thought prompting is not a universal boost, improving DeepSeek-V4-Flash by 32.4 percentage points at L3 but reducing Qwen2.5-7B by 15.6 points, and that L5 graph quality must be scored with partial-credit metrics because node identification and edge extraction diverge sharply. A human pilot under the identical protocol reproduces the same cliff in every language, with L3 at 0% and L4 around 41% ± 3.3 points, supporting a modality-related working-memory bottleneck rather than a language-specific weakness.
Original abstract
Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages has not been systematically studied. We introduce MentalMap, a multilingual diagnostic benchmark with a six-level capability hierarchy (L0-L5) spanning atomic spatial facts to generative world-graph construction, together with four diagnostic axes probing frame of reference, reading-direction bias, reasoning-effort allocation, and hallucination. MentalMap is built from 100 ProcTHOR household scenes, covers eight typologically diverse languages plus a structured-text control, and contains 39 task families across 1,950 evaluation cells. Evaluating thirteen LLMs across scales and model families, we identify a universal L3 reasoning cliff: no model retains even half of its L0 performance on viewpoint reasoning once baseline atomic accuracy exceeds 40%. The cliff persists across languages, scales, and prompting strategies, while structured-output failures and reasoning patterns vary substantially across models. Human evaluation under the identical pure-text protocol reproduces the same failure pattern, suggesting that the bottleneck arises from text-only working memory constraints rather than being specific to current LLM architectures. Our findings reframe pure-text spatial reasoning as a multi-axis world-modeling problem and motivate multimodal and scratchpad-augmented reasoning as future directions.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.