HappyWorld-Bench
AuthorsZhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu
Resources
HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.
Key results
Video World Model track prompts in HappyWorld-Bench
Explicit scenes evaluated for usability, consistency, editing, and expansion
Action-conditioned embodied cases spanning W2–W4
Highest human-preference rating in the video track
Highest stable plate-placement rate among spatial systems
Highest embodied single-action capability score
What the paper found
HappyWorld-Bench evaluates whether AI-generated environments remain coherent, persistent, and responsive under exploration, interaction, and intervention—not merely whether they look realistic. Its six-level W1–W6 framework spans perceptual, interactive, persistent, programmable, scalable, and universal worlds across three tracks containing 1138 video prompts, 300 spatial scenes, and 254 embodied test cases. The benchmark combines HappyWorld-Arena human A/B comparisons and Elo ratings with automated measures based on CLIP, DINOv2, DA3 geometric reprojection, RAFT optical flow, Qwen3-VL, and physics-oriented checks, evaluating 14 video, 9 spatial, and 8 embodied systems. In video, HappyOyster leads human preference at Arena Elo 1263, ahead of Genie 3 at 1206, but extended rollouts still show state drift and weak causal responses. In spatial evaluation, NVIDIA’s Cosmos 3, Marble, HY-World, and GPT-6-Astra demonstrate that visual quality does not guarantee usable geometry: the best placement accuracy is 70.14% and the best edit execution is 73.33%. The embodied track tests single actions, 12-second multi-step sequences, and matched interventions involving altered actions or physical rules. MiniMax-H3 leads embodied capability at W2 score 93.60, while OpenAI’s Sora 2 scores 60.91; ByteDance’s Seedance 2.5, xAI’s Grok Imagine Video 1.5, and Alibaba’s Wan 3.0 are also evaluated. The central finding is a persistent gap between photorealistic generation and reliable action grounding, state persistence, and condition-dependent physical behavior.
Original abstract
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.