ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
AuthorsFan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
Resources
ABot-World-0 aims to make interactive video world models run in real time on a single desktop GPU while preserving controllable, coherent long-horizon environments.
Key results
Parameter count of ABot-World-0.
Strict action-following score on WorldRoamBench.
Maximum FPS at 720P on one NVIDIA RTX 5090.
End-to-end response latency after receiving an action.
Approximate upper memory budget in GiB across optimized low-bit configurations.
What the paper found
Alibaba Group’s AMAP CV Lab presents ABot-World-0, a 5B-parameter action-conditioned video world model designed to turn interactive generation into a local, persistent simulator. It combines AAA games, Unreal Engine and 3D Gaussian Splatting simulations, and internet video, with WorldExplorer adaptively collecting data from model weaknesses and a filtering pipeline using 14 deterministic quality checks plus vision-language assessment. The model uses an 8-key keyboard interface for both camera navigation and third-person character control, while reference-character memory preserves identity. Its main training contribution is a progressive conversion of a bidirectional Wan2.2-based teacher into a causal student through teacher forcing and ODE distillation, followed by LongForcing, which matches long student self-rollouts to an extended-horizon teacher to reduce autoregressive drift. On WorldRoamBench, ABot-World-0 records a strict action accuracy of 0.5266, competitive with systems including Google DeepMind’s Genie 3. For deployment, LightVAE decoding, SageAttention2, Fast-RoPE, bounded KV caching, memory-aware scheduling, and low-bit DiT inference allow 720P streaming at up to 16 FPS on a single NVIDIA RTX 5090, with 1.2 s action-to-first-frame latency and peak VRAM below 19.3 GiB. Extended demonstrations show coherent, controllable rollouts lasting up to 24 hours, although the paper primarily validates stability through sampled checkpoints and a 60-second LongForcing ablation.
Original abstract
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.