Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
AuthorsXin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
Resources
Qwen-Drive-1.0 turns a general vision-language model into an autonomous-driving system that can understand 3D scenes and plan vehicle motion.
Key results
Combined perception, driving vision-language, and general vision-language examples.
Qwen-Drive-1.0-SFT performance on the unified nuScenes validation split.
Semantic occupancy mean IoU on nuScenes.
Average score across the selected driving question-answering benchmarks.
Predictive Driver Model Score after reinforcement learning.
Closed-loop off-road rate, reduced from 24.0% before reinforcement learning.
What the paper found
Qwen-Drive-1.0 is a unified autonomous-driving vision-language model built on Qwen3.5-4B without changing its pretrained architecture. It adds an external bird’s-eye-view perception head for 3D object detection, semantic occupancy, and map segmentation, plus a Planning Expert that uses flow matching to generate 50 future waypoints over 5 seconds. A four-stage recipe combines driving supervision, general vision-language data, and reinforcement learning, with a 1.54M-example Stage 2 mixture designed to improve driving competence while limiting catastrophic forgetting. On nuScenes, the model reaches 43.95 mAP for 3D detection and 60.99 mIoU for occupancy, while its driving VQA average rises to 69.43, outperforming the base Qwen3.5-4B and competing with larger systems such as NVIDIA’s Cosmos models, Gemma4, and Alpamayo. Reinforcement learning raises the NAVSIM Predictive Driver Model Score to 90.7 and produces a 7.91 Rater Feedback Score on the held-out Waymo Open Dataset end-to-end test split. In closed-loop AlpaSim evaluation, reinforcement learning reduces the off-road rate from 24.0% to 12.0%, although progress also falls, highlighting a safety-versus-aggressiveness trade-off. The central contribution is an inspectable 3D interface and trajectory-generation module that extend a general-purpose VLM while largely preserving broad visual reasoning and instruction following.
Original abstract
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.