NTH

HappyWorld-Bench

AuthorsZhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu

September 24, 2026 2 min read
Watch on YouTube
The one-line take

HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.

Key results

1138
Video prompts

Video World Model track prompts in HappyWorld-Bench

300
Spatial scenes

Explicit scenes evaluated for usability, consistency, editing, and expansion

254
Embodied test cases

Action-conditioned embodied cases spanning W2–W4

1263
HappyOyster Arena Elo

Highest human-preference rating in the video track

70.14%
Best spatial placement accuracy

Highest stable plate-placement rate among spatial systems

93.60
MiniMax-H3 W2 score

Highest embodied single-action capability score

What the paper found

HappyWorld-Bench evaluates whether AI-generated environments remain coherent, persistent, and responsive under exploration, interaction, and intervention—not merely whether they look realistic. Its six-level W1–W6 framework spans perceptual, interactive, persistent, programmable, scalable, and universal worlds across three tracks containing 1138 video prompts, 300 spatial scenes, and 254 embodied test cases. The benchmark combines HappyWorld-Arena human A/B comparisons and Elo ratings with automated measures based on CLIP, DINOv2, DA3 geometric reprojection, RAFT optical flow, Qwen3-VL, and physics-oriented checks, evaluating 14 video, 9 spatial, and 8 embodied systems. In video, HappyOyster leads human preference at Arena Elo 1263, ahead of Genie 3 at 1206, but extended rollouts still show state drift and weak causal responses. In spatial evaluation, NVIDIA’s Cosmos 3, Marble, HY-World, and GPT-6-Astra demonstrate that visual quality does not guarantee usable geometry: the best placement accuracy is 70.14% and the best edit execution is 73.33%. The embodied track tests single actions, 12-second multi-step sequences, and matched interventions involving altered actions or physical rules. MiniMax-H3 leads embodied capability at W2 score 93.60, while OpenAI’s Sora 2 scores 60.91; ByteDance’s Seedance 2.5, xAI’s Grok Imagine Video 1.5, and Alibaba’s Wan 3.0 are also evaluated. The central finding is a persistent gap between photorealistic generation and reliable action grounding, state persistence, and condition-dependent physical behavior.

Original abstract

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis