NTH

GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation

AuthorsBoxiang Qiu, Liliang Chen, Yue Liao, Nan Wang, Lintao Wang, Jiayi Luo, Wenzhi Zhao, Shengcong Chen, Di Chen, Ye Li, Chen Gao, Shuicheng Yan, Si Liu, Maoqing Yao, Guanghui Ren

June 3, 2026 2 min read
Watch on YouTube
The one-line take

GE-Sim 2.0 is a faster, more useful robotic world simulator that not only predicts video from actions but also helps train and evaluate manipulation policies in a closed loop.

Key results

2B
Model size

The simulator is built on a 2B-parameter backbone.

23.05
Replay PSNR

Average head-view replay quality across six manipulation tasks.

0.846
Replay SSIM

Average head-view replay structural similarity across six manipulation tasks.

32.28
Replay FID

Average head-view replay Fréchet Inception Distance across six manipulation tasks.

0.81
Closed-loop agreement accuracy

Episode-level agreement between simulated and real robot outcomes in closed-loop evaluation.

What the paper found

GE-Sim 2.0, from AgiBot with collaborators at BUAA, LV-NUS Lab, and TJU, is a closed-loop video world simulator for robotic manipulation that extends the earlier Genie Envisioner framework by retraining a 2B-parameter Cosmos-Predict2-2B-Video2World diffusion transformer on thousands of hours of real robot data, including teleoperation, contact-rich interaction, and on-robot deployment rollouts. Its key novelty is that simulation is no longer view-only: a proprioceptive state expert decodes 16-D joint-angle and gripper state from video latents for dual-arm policies, a VLM-based world judge built on the Robometer formulation turns rollouts into per-frame success rewards, and a DMD2-style acceleration scheme compresses generation to four denoising steps, producing 25 frames in 2.3 seconds on a single NVIDIA H100 and supporting up to 4× frame skipping. On WorldArena, GE-Sim 2.0 ranks first and outperforms specialized robotic world models such as Ctrl-World and DreamDojo, as well as closed-source generators including OpenAI’s Sora and Google DeepMind’s Veo. Across six long-horizon tasks, it improves replay quality to 23.05 dB PSNR, 0.846 SSIM, and 0.145 LPIPS on head view, while reducing FID to 32.28 and FVD to 481.3; in closed-loop evaluation its success agreement with the real robot reaches 0.81 accuracy, and WM-filtered behavior cloning raises downstream real-robot success from 0.417 to 0.567, a 15-point gain.

Original abstract

We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on thousands of hours of real-world robot data spanning teleoperation, contact-rich interaction, and on-robot policy deployment, substantially improving action-following fidelity and trajectory coverage. On top of this foundation, three new modules close the loop from video simulation to policy learning: a state expert that decodes proprioceptive state from video latents to support next-chunk prediction by downstream VLA policies; a world judge that scores generated rollouts against task instructions, yielding machine-verifiable success signals and rewards in place of manual inspection; and an acceleration framework that delivers a 25-frame rollout in 2.3 seconds on a single H100, with up to 4* frame skipping at inference for long-horizon evaluation. GE-Sim 2.0 tops the public WorldArena leaderboard at only 2B parameters, outperforming both dedicated robotic world models and closed-source general video generators, and policies trained against its rollouts and rewards translate into measurable real-world gains, establishing GE-Sim 2.0 as a practical platform for scalable evaluation and closed-loop learning of manipulation policies.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis