NTH

Hallucination in World Models is Predictable and Preventable

AuthorsNicklas Hansen, Xiaolong Wang

July 3, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that hallucinations in world models are not random: they happen where training data is sparse, and the authors use that insight to detect, reduce, and fine-tune around them.

Key results

427
MMBench2 hours

visual world-modeling dataset scale

210
MMBench2 tasks

number of tasks spanning 10 domains

350M
World model size

Dreamer 4 world model parameter count

0.80
Predictor correlation

magnitude of Spearman correlation with rollout ΔPSNR

0.88
Rollout ΔPSNR gain

improvement from coverage-aware training on held-out trajectories

50
Unseen-task trajectories

real trajectories per task needed for adaptation

What the paper found

Nicklas Hansen and Xiaolong Wang at UC San Diego show that hallucination in modern world models is not mainly an architectural flaw but a data-coverage problem: it concentrates in low-coverage regions of state-action space and can be predicted from internal signals. Using their new MMBench2 benchmark, a 427-hour, 210-task corpus with ground-truth actions, rewards, and live simulators, they train a 350M-parameter Dreamer 4 world model and separate hallucination into three failure modes: perceptual errors in the tokenizer, action-marginalized transitions in the dynamics model, and scene-diverging rollouts. They then introduce three label-free predictors—the tokenizer round-trip residual, flow instability, and inter-seed variance—that correlate strongly with open-loop rollout error, with Spearman ρ around −0.80 across all three. For mitigation, they reweight sampling to be uniform across tasks rather than frames, which improves rollout fidelity by 0.88 dB and lowers all three normalized hallucination signals; they also use the predictors as curiosity rewards for online data collection, where only 50 real trajectories per task adapt the pretrained model to entirely unseen environments. On 10 unseen tasks, closed-loop MPC performance rises from 0.276 to 0.325 with curiosity-driven data, compared with 0.118 for a random policy and 0.362 for expert or human data, showing that predictive hallucination signals can be turned directly into data-efficient control improvements.

Original abstract

Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation. To test this, we introduce MMBench2, a 427-hour, 210-task dataset for visual world modeling with ground-truth actions, rewards, and live simulators, and train a 350M-parameter world model on it. We identify three distinct hallucination modes: perceptual, action-marginalized, and scene-diverging -- each anchored to a different stage of the pipeline, and develop three signals that accurately predict where the model will fail. To close coverage gaps at training time, we develop a coverage-aware sampling technique; to close them online, our hallucination predictors serve as curiosity rewards for targeted data collection, yielding a data-efficient finetuning recipe that adapts the pretrained world model to entirely unseen environments with as few as 50 real environment trajectories. Overall, our findings reveal that hallucination in world models is inherently a data coverage issue, and that the same signals used to detect it can also be used for mitigation. An interactive web version of our paper is available at https://www.nicklashansen.com/mmbench2

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis