Orca: The World is in Your Mind
AuthorsYihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen, Runze Xiao, Huaihai Lyu, Senwei Xie, Euan Liu, Klara Tian, Tianfeng Long, Yichi Zhang, Zhengliang Cai, Ruike Chen, Jifan Zhao, Ruochuan Shi, Zihan Tang, Jing Lyu, Wenxing Tan, Ningbo Zhang, Yangtao Hu, Yuming Gao, Xiansheng Chen, Junkai Zhao, Congsheng Xu, Boan Zhu, Ziqi Wang, Yupu Feng, Qiongqiong Zhang, Yingli Zhao, Yulong Ao, Shaoxuan Xie, You Liu, Guocai Yao, Leiduo Zhang, Xiaodan Liu, Yunyan Zhang, Yance Jiao, Xinyan Yang, Jiaxing Wei, Xu Liu, Tengfei Pan, Shaokai Nie, Chunlei Men, Sen Cui, Xiaojie Jin, Hongyang Li, Jianlan Luo, Yao Mu, Yunchao Wei, Jun Yan, Hang Zhao, Xiaolong Zheng, Jiaming Li, Yonghua Lin, Tiejun Huang, Zhongyuan Wang, Pengwei Wang
Resources
Orca is a general world model that learns a shared latent representation from video, language, and other signals so it can better predict, describe, and act in the world.
Key results
Approximate video hours used in pre-training
Event annotations in the pre-training inventory
General VQA data in the pre-training inventory
Average zero-shot text generation score on MVBench, TemporalBench, 3DSRBench, and SWITCH
Average score on PRICE-V0.1 with the 4B+2B readout
Overall out-of-distribution rule-based action score
What the paper found
Orca, from the Beijing Academy of Artificial Intelligence, reframes multimodal foundation modeling around next-state prediction rather than next-token, next-frame, or next-action prediction, using a frozen Qwen3.5 VLM backbone to learn a unified world latent space from vision and language. The model is trained with two complementary paradigms: unconscious learning over dense video transitions and conscious learning over language-conditioned event transitions plus VQA supervision, using a loss mix of 0.1 for observation-only transition, 0.5 for event-conditioned transition, and 0.4 for VQA. Pre-training uses 12.5K hours of video, 160M event annotations, and an additional 11.5M VQA data points, though only one-tenth of the inventory is used in this version. With a frozen encoder and lightweight decoders, Orca scales from 0.8B to 4B parameters and improves zero-shot text generation on MVBench, TemporalBench, 3DSRBench, and SWITCH, reaching 51.8 average at 4B; it also tops the real-world PRICE-V0.1 image-prediction benchmark at 59.8 average with a 4B+2B readout. In embodied action generation on five dual-arm robot tasks, Orca achieves 32.4 overall rule-based score in out-of-distribution settings, outperforming Qwen3.5 and approaching 𝜋0.5. The engineering stack matters too: FlagScale plus FSDP2, chunked cross-entropy, activation recomputation, and communication prefetching raise throughput from 0.66 to 2.91 samples/sec/GPU, a 4.4× speedup over StarVLA.
Original abstract
We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multimodal readout interfaces. Rather than optimizing isolated next-token, next-frame, or next-action prediction, we are centered on Next-State-Prediction modeling, offering a unified state-transition modeling route toward understanding, predicting, and acting upon the world. Orca learns through two complementary paradigms: unconscious learning captures dense natural state transitions from continuous videos, and conscious learning models sparse meaningful state transitions by language-described events and VQA supervision. For pre-training, we construct a large-scale world-learning inventory data, including 125K hours of video data and 160M event annotations. After pre-training, Orca learns a unified world latent space. To examine whether the learned latent supports downstream, we evaluate it by three representative downstream readouts: text generation, image prediction, and embodied action generation. Orca's backbone is frozen, and only the lightweight modality-specific decoders are trainable. Experiments show the scalability of the proposed paradigm and verify that stronger world latent enables stronger downstream readouts. Orca outperforms similar-sized specialized baselines. These results show that Orca, as a general world foundation model, presents a promising approach to understanding, predicting, and acting upon the world. Finally, we discuss the current limitations, aiming to provide useful insights and inspiration for the community.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.