NTH

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

AuthorsXin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai

September 5, 2026 2 min read
Watch on YouTube
The one-line take

Qwen-Drive-1.0 turns a general vision-language model into an autonomous-driving system that can understand 3D scenes and plan vehicle motion.

Key results

1.54M
Stage 2 training mixture

Combined perception, driving vision-language, and general vision-language examples.

43.95
nuScenes 3D detection mAP

Qwen-Drive-1.0-SFT performance on the unified nuScenes validation split.

60.99
nuScenes occupancy mIoU

Semantic occupancy mean IoU on nuScenes.

69.43
Driving VQA average

Average score across the selected driving question-answering benchmarks.

90.7
NAVSIM PDMS

Predictive Driver Model Score after reinforcement learning.

12.0%
AlpaSim off-road rate after reinforcement learning

Closed-loop off-road rate, reduced from 24.0% before reinforcement learning.

What the paper found

Qwen-Drive-1.0 is a unified autonomous-driving vision-language model built on Qwen3.5-4B without changing its pretrained architecture. It adds an external bird’s-eye-view perception head for 3D object detection, semantic occupancy, and map segmentation, plus a Planning Expert that uses flow matching to generate 50 future waypoints over 5 seconds. A four-stage recipe combines driving supervision, general vision-language data, and reinforcement learning, with a 1.54M-example Stage 2 mixture designed to improve driving competence while limiting catastrophic forgetting. On nuScenes, the model reaches 43.95 mAP for 3D detection and 60.99 mIoU for occupancy, while its driving VQA average rises to 69.43, outperforming the base Qwen3.5-4B and competing with larger systems such as NVIDIA’s Cosmos models, Gemma4, and Alpamayo. Reinforcement learning raises the NAVSIM Predictive Driver Model Score to 90.7 and produces a 7.91 Rater Feedback Score on the held-out Waymo Open Dataset end-to-end test split. In closed-loop AlpaSim evaluation, reinforcement learning reduces the off-road rate from 24.0% to 12.0%, although progress also falls, highlighting a safety-versus-aggressiveness trade-off. The central contribution is an inspectable 3D interface and trajectory-generation module that extend a general-purpose VLM while largely preserving broad visual reasoning and instruction following.

Original abstract

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis