DriveZero: End-to-End Driving Beyond Human Demonstrations
AuthorsHao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang
Resources
DriveZero teaches autonomous cars to drive beyond human examples by combining foundation-model perception with reinforcement learning in simulated interactive worlds.
Key results
Privileged action model trained with PPO from scratch
Mixed-agent worlds simulated in parallel across 96 GPUs
DriveRL with value-guided test-time action search across six settings
DriveZero-Scale camera-only performance without human trajectory supervision
State-of-the-art performance after scaling with SimScale data
Zero-shot closed-loop performance
What the paper found
Xiaomi EV’s DriveZero addresses a core limitation of imitation-based autonomous driving: human logs provide only one demonstrated future and rarely contain recovery behavior or valid alternatives. Its solution separates perception from action. DriveRL is a 5.70M-parameter privileged teacher trained from scratch with PPO in mixed-agent, closed-loop worlds generated from nuPlan logs, scaling to 196,608 worlds in parallel; value-guided test-time action search raises its mean nuPlan score to 93.57 across six reactive and non-reactive settings, exceeding the Log-Replay expert. DriveVFM builds the visual backbone by distilling DINOv3, SigLIP2, SAM, and Depth Anything V2 through feature matching, using web and driving imagery without task-specific perception labels. DriveZero then initializes from DriveVFM and distills goal-conditioned DriveRL rollouts into a camera-only Transformer planner with multiple trajectory proposals, winner-takes-all supervision, and proposal scoring. Because the teacher can be queried with augmented navigation goals, the student receives diverse, goal-consistent trajectories rather than a single human future. Without human trajectory supervision, DriveZero-Scale reaches 95.3 PDMS on NAVSIMv1 navtest, surpassing the human-driver score of 94.8, achieves 57.1 EPDMS on NAVSIMv2 navhard, and obtains 46.6 average HD-Score zero-shot on the closed-loop HUGSIM benchmark. The central contribution is a scalable perception-action decomposition that transfers behavior learned through reinforcement and closed-loop feedback into a deployable visual policy.
Original abstract
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.