NTH

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

AuthorsQuanfu Yu, Xian Wu, Hao Xu, Liulong Ma

July 30, 2026 2 min read
Watch on YouTube
The one-line take

HyWorldVLA combines pixel-grounded and latent world modeling to make autonomous-driving vision-language-action systems more robust to noisy future predictions.

Key results

90.59
NAVSIM v1 PDMS

Overall score achieved with a single front-view camera.

89.71
NAVSIM v2 EPDMS

Extended benchmark score across compliance, safety, progress, lane keeping, and comfort metrics.

655
Noise robustness test cases

Rain- and fog-corrupted cases used for robustness evaluation.

86.87
Noise robustness PDMS

HyWorldVLA score on corrupted cases, versus 61.18 for DriveVLA-W0.

87.50
Pure pixel-world ablation PDMS

Score after removing latent representation capacity.

What the paper found

Researchers at BYD Company Limited introduce HyWorldVLA, a vision-language-action model designed to reconcile the detailed reasoning of pixel-based world models with the noise robustness of latent models. Built on Emu3, the system first trains a text-guided VideoVAEPlus to compress future driving video into spatiotemporal latents, using Flan-T5 cross-attention to improve reconstruction. During pre-training, HyWorldVLA jointly predicts discrete visual tokens, FAST action tokens, language tokens, and future VAE latents while also reconstructing video frames; during co-fine-tuning, it uses only predicted latent features, historical motion, and commands to condition a query-based action expert that generates trajectories. With a single front-view camera, the model reaches a NAVSIM v1 PDMS of 90.59 and a NAVSIM v2 EPDMS of 89.71, surpassing pixel-based and latent-based baselines. The paper also introduces a rain-and-fog robustness set containing 655 cases: HyWorldVLA scores 86.87, compared with 61.18 for DriveVLA-W0. Ablations show that removing pixel-level supervision reduces PDMS to 87.50, while removing latent modeling yields 89.91, supporting the central claim that both forms of supervision are complementary. The study positions hybrid world modeling as a practical route toward more interpretable and stable end-to-end autonomous driving.

Original abstract

Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis