NTH

Cosmos 3: Omnimodal World Models for Physical AI

AuthorsAditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilat...

June 17, 2026 2 min read
Watch on YouTube
The one-line take

Cosmos 3 is a unified multimodal world model that can understand and generate text, images, video, audio, and actions, aiming to become a general backbone for embodied AI.

Key results

24.2M
Reasoner samples

Total reasoner training corpus across pre-training and supervised fine-tuning.

31.05T
Cosmos3-Nano pre-training tokens

Tokens used to pre-train Cosmos3-Nano.

17.86T
Cosmos3-Super pre-training tokens

Tokens used to pre-train Cosmos3-Super.

91.36
UniGenBench score

Cosmos3-Super-Text2Image score on UniGenBench.

89.3
Cosmos-HUE T2V score

Cosmos3-Super score on Cosmos-HUE text-to-video.

39.7%
RoboLab success

Cosmos3-Nano-Policy-DROID success rate under specific instructions on RoboLab.

What the paper found

NVIDIA’s Cosmos 3 is an omnimodal world-model family that unifies language, image, video, audio, and action in a single Mixture-of-Transformers architecture, with separate autoregressive and diffusion towers sharing multimodal attention and 3D MRoPE plus absolute temporal modulation. The paper’s core claim is that this unified design can replace fragmented VLM, video-generation, world-simulation, and policy pipelines, and it scales from a 4B Edge model to 16B Nano and 64B Super variants. Cosmos 3 is trained in staged curricula on 24.2M reasoner samples, 31.05T tokens for Cosmos3-Nano pre-training, and 17.86T tokens for Cosmos3-Super pre-training, then post-trained into specialists such as Cosmos3-Super-Text2Image, Cosmos3-Super-Image2Video, and Cosmos3-Nano-Policy-DROID. On evaluation, Cosmos3-Super reaches 91.36 on UniGenBench, 89.3 on Cosmos-HUE text-to-video, 89.6 on Cosmos-HUE image-to-video, and 63.4 on Physics-IQ video-to-video with WMReward + best-of-N, while Cosmos3-Nano-Policy-DROID tops RoboLab with 39.7% success under specific instructions and ranks first on RoboArena. The infrastructure story is equally central: NVIDIA reports SILA data processing cut job startup latency from 30–60 minutes to about 5 minutes and delivered a 10× throughput increase, while training optimizations such as rank-synchronous stream selection, selective activation checkpointing, and torch.compile materially improved throughput on GB200 systems.

Original abstract

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 https://openmdw.ai/license/1-1/ License at https://github.com/nvidia/cosmos}{github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3 . The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3 .

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis