NTH

Orca: The World is in Your Mind

AuthorsYihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen, Runze Xiao, Huaihai Lyu, Senwei Xie, Euan Liu, Klara Tian, Tianfeng Long, Yichi Zhang, Zhengliang Cai, Ruike Chen, Jifan Zhao, Ruochuan Shi, Zihan Tang, Jing Lyu, Wenxing Tan, Ningbo Zhang, Yangtao Hu, Yuming Gao, Xiansheng Chen, Junkai Zhao, Congsheng Xu, Boan Zhu, Ziqi Wang, Yupu Feng, Qiongqiong Zhang, Yingli Zhao, Yulong Ao, Shaoxuan Xie, You Liu, Guocai Yao, Leiduo Zhang, Xiaodan Liu, Yunyan Zhang, Yance Jiao, Xinyan Yang, Jiaxing Wei, Xu Liu, Tengfei Pan, Shaokai Nie, Chunlei Men, Sen Cui, Xiaojie Jin, Hongyang Li, Jianlan Luo, Yao Mu, Yunchao Wei, Jun Yan, Hang Zhao, Xiaolong Zheng, Jiaming Li, Yonghua Lin, Tiejun Huang, Zhongyuan Wang, Pengwei Wang

July 3, 2026 2 min read
Watch on YouTube
The one-line take

Orca is a general world model that learns a shared latent representation from video, language, and other signals so it can better predict, describe, and act in the world.

Key results

12.5K
video_hours

Approximate video hours used in pre-training

160M
event_annotations

Event annotations in the pre-training inventory

11.5M
vqa_data

General VQA data in the pre-training inventory

51.8
text_generation_avg_4B

Average zero-shot text generation score on MVBench, TemporalBench, 3DSRBench, and SWITCH

59.8
price_v0_1_avg_4B

Average score on PRICE-V0.1 with the 4B+2B readout

32.4
action_rule_based_overall

Overall out-of-distribution rule-based action score

What the paper found

Orca, from the Beijing Academy of Artificial Intelligence, reframes multimodal foundation modeling around next-state prediction rather than next-token, next-frame, or next-action prediction, using a frozen Qwen3.5 VLM backbone to learn a unified world latent space from vision and language. The model is trained with two complementary paradigms: unconscious learning over dense video transitions and conscious learning over language-conditioned event transitions plus VQA supervision, using a loss mix of 0.1 for observation-only transition, 0.5 for event-conditioned transition, and 0.4 for VQA. Pre-training uses 12.5K hours of video, 160M event annotations, and an additional 11.5M VQA data points, though only one-tenth of the inventory is used in this version. With a frozen encoder and lightweight decoders, Orca scales from 0.8B to 4B parameters and improves zero-shot text generation on MVBench, TemporalBench, 3DSRBench, and SWITCH, reaching 51.8 average at 4B; it also tops the real-world PRICE-V0.1 image-prediction benchmark at 59.8 average with a 4B+2B readout. In embodied action generation on five dual-arm robot tasks, Orca achieves 32.4 overall rule-based score in out-of-distribution settings, outperforming Qwen3.5 and approaching 𝜋0.5. The engineering stack matters too: FlagScale plus FSDP2, chunked cross-entropy, activation recomputation, and communication prefetching raise throughput from 0.66 to 2.91 samples/sec/GPU, a 4.4× speedup over StarVLA.

Original abstract

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multimodal readout interfaces. Rather than optimizing isolated next-token, next-frame, or next-action prediction, we are centered on Next-State-Prediction modeling, offering a unified state-transition modeling route toward understanding, predicting, and acting upon the world. Orca learns through two complementary paradigms: unconscious learning captures dense natural state transitions from continuous videos, and conscious learning models sparse meaningful state transitions by language-described events and VQA supervision. For pre-training, we construct a large-scale world-learning inventory data, including 125K hours of video data and 160M event annotations. After pre-training, Orca learns a unified world latent space. To examine whether the learned latent supports downstream, we evaluate it by three representative downstream readouts: text generation, image prediction, and embodied action generation. Orca's backbone is frozen, and only the lightweight modality-specific decoders are trainable. Experiments show the scalability of the proposed paradigm and verify that stronger world latent enables stronger downstream readouts. Orca outperforms similar-sized specialized baselines. These results show that Orca, as a general world foundation model, presents a promising approach to understanding, predicting, and acting upon the world. Finally, we discuss the current limitations, aiming to provide useful insights and inspiration for the community.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis