GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
AuthorsGigaWorld Team, Angyuan Ma, Boyuan Wang, Bohan Li, Chaojun Ni, Guo Li, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jingyu Liu, Jiwen Lu, Qiuping Deng, Tingdong Yu, Xuancheng Xu, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Xiaofeng Wang, Xiaoyu Tian, Yang Wang, Yifan Chang, Yukun Zhou, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
Resources
This paper shows how world models can serve as practical stand-ins for expensive robot testing, and introduces a benchmark and model designed to make that evaluation much more reliable.
Key results
WMBench corpus size across 8 tasks
world-model rollout segments analyzed for evaluator quality
mixed data used to train GigaWorld-1
best overall evaluator-relevant score
relative gain in average score
Qwen3-VL evaluator vs human WMES labels
What the paper found
GigaWorld-1 from GigaAI and Tsinghua University reframes world modeling as a first-class problem for robot policy evaluation, where the goal is not photorealistic video synthesis but agreement with real-robot success and failure under closed-loop rollout. The paper introduces WMBench, a benchmark built from 2,989 paired trajectories across 8 manipulation tasks, then studies 7 world models, 4 action representations, and 324,000+ simulated rollouts to identify what actually predicts evaluator quality. The strongest signal is not appearance stability but long-horizon, action-faithful consistency: Subject Consistency reaches 0.88 correlation with the World Model as Evaluator Score, Perspectivity 0.86, and Instruction Following 0.84, while Background Consistency and Photometric Consistency are negatively correlated at -0.45 and -0.42. On the modeling side, the best action interface is a pixel-aligned control map, which raises Trajectory Accuracy to 0.3528, and hierarchical memory dramatically improves 40-second rollout quality, with GigaWorld-1+Mem pushing PSNR to 19.82 and FID down to 40.58 from 14.05 and 142.84 in the memory-free baseline. After training on about 12,980 hours of mixed physical, robot, egocentric, and Giga-collected data, GigaWorld-1-Plus reaches 0.6834 average score, outperforming Cosmos-Predict2.5 by 11.6% and Wan 2.2 5B by 14.9%, while a LoRA-tuned Qwen3-VL evaluator achieves 87.80% exact agreement and 0.7349 quadratic weighted kappa against human WMES labels.
Original abstract
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in \textit{GigaWorld-1}, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.