NTH

DriveZero: End-to-End Driving Beyond Human Demonstrations

AuthorsHao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang

September 13, 2026 3 min read
Watch on YouTube
The one-line take

DriveZero teaches autonomous cars to drive beyond human examples by combining foundation-model perception with reinforcement learning in simulated interactive worlds.

Key results

5.70M
DriveRL teacher size

Privileged action model trained with PPO from scratch

196,608
Parallel training worlds

Mixed-agent worlds simulated in parallel across 96 GPUs

93.57
nuPlan mean score

DriveRL with value-guided test-time action search across six settings

95.3
NAVSIMv1 navtest PDMS

DriveZero-Scale camera-only performance without human trajectory supervision

57.1
NAVSIMv2 navhard EPDMS

State-of-the-art performance after scaling with SimScale data

46.6
HUGSIM average HD-Score

Zero-shot closed-loop performance

What the paper found

Xiaomi EV’s DriveZero addresses a core limitation of imitation-based autonomous driving: human logs provide only one demonstrated future and rarely contain recovery behavior or valid alternatives. Its solution separates perception from action. DriveRL is a 5.70M-parameter privileged teacher trained from scratch with PPO in mixed-agent, closed-loop worlds generated from nuPlan logs, scaling to 196,608 worlds in parallel; value-guided test-time action search raises its mean nuPlan score to 93.57 across six reactive and non-reactive settings, exceeding the Log-Replay expert. DriveVFM builds the visual backbone by distilling DINOv3, SigLIP2, SAM, and Depth Anything V2 through feature matching, using web and driving imagery without task-specific perception labels. DriveZero then initializes from DriveVFM and distills goal-conditioned DriveRL rollouts into a camera-only Transformer planner with multiple trajectory proposals, winner-takes-all supervision, and proposal scoring. Because the teacher can be queried with augmented navigation goals, the student receives diverse, goal-consistent trajectories rather than a single human future. Without human trajectory supervision, DriveZero-Scale reaches 95.3 PDMS on NAVSIMv1 navtest, surpassing the human-driver score of 94.8, achieves 57.1 EPDMS on NAVSIMv2 navhard, and obtains 46.6 average HD-Score zero-shot on the closed-loop HUGSIM benchmark. The central contribution is a scalable perception-action decomposition that transfers behavior learned through reinforcement and closed-loop feedback into a deployable visual policy.

Original abstract

Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis