HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving
AuthorsQuanfu Yu, Xian Wu, Hao Xu, Liulong Ma
Resources
HyWorldVLA combines pixel-grounded and latent world modeling to make autonomous-driving vision-language-action systems more robust to noisy future predictions.
Key results
Overall score achieved with a single front-view camera.
Extended benchmark score across compliance, safety, progress, lane keeping, and comfort metrics.
Rain- and fog-corrupted cases used for robustness evaluation.
HyWorldVLA score on corrupted cases, versus 61.18 for DriveVLA-W0.
Score after removing latent representation capacity.
What the paper found
Researchers at BYD Company Limited introduce HyWorldVLA, a vision-language-action model designed to reconcile the detailed reasoning of pixel-based world models with the noise robustness of latent models. Built on Emu3, the system first trains a text-guided VideoVAEPlus to compress future driving video into spatiotemporal latents, using Flan-T5 cross-attention to improve reconstruction. During pre-training, HyWorldVLA jointly predicts discrete visual tokens, FAST action tokens, language tokens, and future VAE latents while also reconstructing video frames; during co-fine-tuning, it uses only predicted latent features, historical motion, and commands to condition a query-based action expert that generates trajectories. With a single front-view camera, the model reaches a NAVSIM v1 PDMS of 90.59 and a NAVSIM v2 EPDMS of 89.71, surpassing pixel-based and latent-based baselines. The paper also introduces a rain-and-fog robustness set containing 655 cases: HyWorldVLA scores 86.87, compared with 61.18 for DriveVLA-W0. Ablations show that removing pixel-level supervision reduces PDMS to 87.50, while removing latent modeling yields 89.91, supporting the central claim that both forms of supervision are complementary. The study positions hybrid world modeling as a practical route toward more interpretable and stable end-to-end autonomous driving.
Original abstract
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.