NTH

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

AuthorsZhikai Pan, Chih-Ting Liao, Chunrui Liu, Xi Xiao, Yitong Qiao, Chunlei Meng, Zhangquan Chen, Xin Cao

June 3, 2026 2 min read
Watch on YouTube
The one-line take

This paper tests whether LLMs can build spatial mental maps from text across languages, and finds a consistent reasoning cliff that suggests current models struggle to form robust world models.

Key results

100
ProcTHOR houses

MENTAL MAP is grounded in 100 ProcTHOR household scenes.

39
task families

MENTAL MAP comprises 39 task families across the benchmark.

1950
evaluation cells

The benchmark is evaluated across 1,950 cells.

0%
L3 cliff human pilot

The human pilot reproduces the cliff with L3 equal to 0% in all eight languages.

41%
L4 human pilot

In the human pilot, L4 is reported as around 41% with 3.3pp standard deviation across languages.

What the paper found

In “Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning,” Zhikai Pan, Chih-Ting Liao, and colleagues from the University of New South Wales introduce MENTAL MAP, a pure-text benchmark built on 100 ProcTHOR household scenes to test whether LLMs form internal spatial world models or just exploit co-occurrence. The benchmark spans 39 task families across 1,950 evaluation cells in eight typologically diverse languages—English, Mandarin, Japanese, Korean, Spanish, Arabic, Thai, and German—plus a structured-text control, and it separates static comprehension from active spatial reasoning through a six-level staircase from L0 atomic facts to L5 generative world-graph output. Across 13 models, including OpenAI’s GPT-4o, Google’s Gemini-2.5-Flash, DeepSeek-V4-Flash, and Qwen family models, the paper finds a universal “L3 cliff”: performance collapses at viewpoint transformation, with all cliff-diagnostic model-language cells scoring below half of their L0 atomic baseline and most below a quarter. The study also shows that chain-of-thought prompting is not a universal boost, improving DeepSeek-V4-Flash by 32.4 percentage points at L3 but reducing Qwen2.5-7B by 15.6 points, and that L5 graph quality must be scored with partial-credit metrics because node identification and edge extraction diverge sharply. A human pilot under the identical protocol reproduces the same cliff in every language, with L3 at 0% and L4 around 41% ± 3.3 points, supporting a modality-related working-memory bottleneck rather than a language-specific weakness.

Original abstract

Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages has not been systematically studied. We introduce MentalMap, a multilingual diagnostic benchmark with a six-level capability hierarchy (L0-L5) spanning atomic spatial facts to generative world-graph construction, together with four diagnostic axes probing frame of reference, reading-direction bias, reasoning-effort allocation, and hallucination. MentalMap is built from 100 ProcTHOR household scenes, covers eight typologically diverse languages plus a structured-text control, and contains 39 task families across 1,950 evaluation cells. Evaluating thirteen LLMs across scales and model families, we identify a universal L3 reasoning cliff: no model retains even half of its L0 performance on viewpoint reasoning once baseline atomic accuracy exceeds 40%. The cliff persists across languages, scales, and prompting strategies, while structured-output failures and reasoning patterns vary substantially across models. Human evaluation under the identical pure-text protocol reproduces the same failure pattern, suggesting that the bottleneck arises from text-only working memory constraints rather than being specific to current LLM architectures. Our findings reframe pure-text spatial reasoning as a multi-axis world-modeling problem and motivate multimodal and scratchpad-augmented reasoning as future directions.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis