NTH

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AuthorsAgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai Wang, Xindong He, Jinyuan Feng, Ziyu Xiong, Linqing Zhong, Yifei Wei, Feng Han, Long Zhang, Da Huang, Nanshu Zhao, Chenghao Yin, Mo Wu, Zhaodong Yan, Kongtao Hu, Yuxiang Yan, Aogelijiang Niyazi, Yu Fang, Jia Zeng, Lizhu Meng, Daizhen Lv, Haoyu Cao, Zhiwen Hou, Lianjin Ye, Yuehan Niu, Zhikai Cai, Xuan Hu, Hui Min, Xiongfeng Cai, Yue Liao, Jing Wu, Soujanya Poria, Ye Li, Sanping Zhou, Maoqing Yao

September 15, 2026 3 min read
Watch on YouTube
The one-line take

GE-Act 2.0 trains a scalable robot world-action model from scratch and shows that more diverse manipulation data can substantially improve zero-shot control across tasks, embodiments, and environments.

Key results

64×
CoAE spatial downsampling

The control-oriented autoencoder compresses each frame with 64× spatial downsampling.

24
CoAE tokens per frame

A 256×384 frame becomes 24 visual latent tokens.

44.1%
G1-OP success at 30,000 hours

Zero-shot OOD success rises from 17.1% at 300 hours to 44.1% at 30,000 hours.

31.1%
G2-90D success at 30,000 hours

Zero-shot OOD success reaches 31.1% on the data-scarce G2-90D embodiment.

17.7
G2-90D scaling gain

Scaling produces a 17.7-point gain on G2-90D.

37.5%
KASO four-object pick success

KASO achieves 37.5% macro-average pick success versus 22.5% for E2E with retained pretraining losses.

What the paper found

GE-Act 2.0 is a world-action model trained from scratch for robotic manipulation rather than adapted from a pretrained video generator such as NVIDIA Cosmos. Its control-oriented autoencoder compresses each 256×384 frame with 64× spatial downsampling into 24 visual tokens, while multi-teacher alignment uses SigLIP 2, V-JEPA 2.1, and DINOv3. A single-step visual planner based on conditional MeanFlow predicts complete future visual states in one differentiable pass, allowing its inverse dynamics model to pretrain independently on instruction-free robot trajectories, including failures and rollouts. The components are connected with Knowledge-Aligned Selective Optimization, or KASO, which samples futures and retains only those behaviorally compatible with the recorded action, addressing the validity gap caused by multimodal demonstrations. Evaluated without task-specific fine-tuning on 100 real-robot tasks spanning 20 skill groups, scaling co-training from 300 hours to 30,000 hours raises zero-shot OOD success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D, a 17.7-point gain for the sparsely represented embodiment. In a controlled ablation, KASO reaches 37.5% four-object pick success versus 22.5% for E2E with retained pretraining losses. The model uses frozen Qwen3.5-2B for scene-grounded language conditioning and compares favorably with models including NVIDIA GR00T N1.7 and Physical Intelligence’s π0.5 on selected simulation robustness benchmarks.

Original abstract

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis