NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
AuthorsNVIDIA, :, Aarti Basant, Amlan Kar, Despoina Paschalidou, Fangyin Wei, Francesco Ferroni, Guillermo Garcia Cobo, Haithem Turki, Huan Ling, Jaewoo Seo, James Lucas, Jay Zhangjie Wu, Jialiang Wang, Jonathan Lorraine, Jun Gao, Kai He, Katarina Tothova, Kevin Xie, Michał Tyszkiewicz, Qi Wu, Riccardo de Lutio, Ruilong Li, Sanja Fidler, Seung Wook Kim, Tianchang Shen, Tianshi Cao, Tobias Pfaff, William Lew, Xindi Wu, Xuanchi Ren, Yifan Lu, Yuxuan Zhang, Zan Gojcic, Zian Wang
Resources
OmniDreams is a real-time AI driving simulator that generates photorealistic future scenes from actions, aiming to make autonomous vehicle testing safer, broader, and more realistic.
Key results
On the held-out RDS-HQ-1M evaluation, the causal student model achieved FVD 31.7 before distillation.
The final distilled OmniDreams model improved FVD to 24.8 on the held-out RDS-HQ-1M evaluation.
The final distilled model reached BEVFormer LET-AP of 0.400 on the held-out RDS-HQ-1M evaluation.
In closed-loop evaluation with Alpamayo 1.5, the baseline collision rate was 6.9% before the OmniDreams-derived WAM fine-tuning improvement.
In closed-loop evaluation with Alpamayo 1.5, the OmniDreams-derived World-Action Model reduced collision rate to 4.2%.
What the paper found
NVIDIA’s OmniDreams paper introduces a real-time generative world model for closed-loop autonomous vehicle simulation, built by mid- and post-training the Cosmos-Predict 2.5 diffusion backbone on 21,544 hours of driving data from RDS and RDS-HQ-1M across 15 countries. Unlike reconstruction-based simulators such as NVIDIA NuRec, OmniDreams is action-conditioned and autoregressive: it generates the next camera observations from a first-frame seed, a structured world-scenario map with HD lanes and 3D tracked agents, a text prompt for weather and lighting, and a streaming KV cache for long-horizon memory. The system uses a causal diffusion transformer, Diffusion Forcing for autoregressive adaptation, and Self Forcing plus Distribution Matching Distillation to reduce exposure bias and make few-step rollout practical. NVIDIA reports 68 FPS for the single-view 2B model on one GB300 and 105 FPS per camera for the 4-view version on a 16-GPU GB300 cabinet at 704×1280. On the held-out RDS-HQ-1M evaluation, the final distilled model improved FVD from 31.7 to 24.8 versus the causal student and also raised BEVFormer LET-AP from 0.221 to 0.400. In closed-loop tests inside AlpaSim with the Alpamayo 1.5 policy, OmniDreams preserved the policy ranking seen under NuRec and reduced collision rate from 6.9% to 4.2% when the OmniDreams backbone was fine-tuned into a World-Action Model, using about 2B parameters versus roughly 10B for Alpamayo 1.5. The paper also shows that OmniDreams can function as a diffusion fixer, correcting reconstruction artifacts while preserving scene geometry and driving-relevant structure.
Original abstract
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations. While recent reconstruction-based neural simulators offer photorealism, they are fundamentally constrained by their initial captured data and struggle to generalize to highly dynamic or novel scenes. To overcome these limitations, we introduce OmniDreams, a foundation generative world model mid- and post-trained from the Cosmos diffusion model to autoregressively generate action-conditioned videos in real time. By leveraging the rich visual priors of Cosmos and mid- and post-training on 21k hours of driving scenarios, OmniDreams synthesizes complex, unobserved phenomena that are hard for traditional simulators to capture, such as extreme weather and unpredictable dynamic agent behaviors. Crucially, it autoregressively conditions its photorealistic sensor generation on past frames, the current simulator state, and immediate driving actions. Deployed in a closed-loop system with the Alpamayo 1 policy model and AlpaSim orchestrator, OmniDreams acts as a highly responsive, reactive environment, providing a scalable and comprehensive solution for training and evaluating next-generation autonomous driving policies. We additionally show preliminary results indicating that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters. These results highlight the potential for a real-time world model like OmniDreams to also serve as a backbone for policy architectures.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.