NTH

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

AuthorsRichard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert, Nils Rack, Kaiwen Wang, Fabian Konstantinidis, Julian Truetsch, Carlos Fernandez, Annika Bätz, Kevin Rösch, Marlon Steiner, Willi Poh, Yinzhe Shen, Royden Wagner, Felix Hauser, Dominik Strutz, Jaime Villa, Gleb Stepanov, Holger Caesar, Ömer Şahin Taş, Frank Bieder, Jan-Hendrik Pauls, Christoph Stiller

July 15, 2026 2 min read
Watch on YouTube
The one-line take

KITScenes Multimodal is a high-fidelity autonomous driving dataset that pairs rich sensors and detailed maps with new benchmarks for mapping, depth, view synthesis, and end-to-end driving.

Key results

72.5 Mpx
Camera resolution per frame

Combined synchronized global-shutter camera resolution.

409.2 m
Maximum lidar range

Measured maximum range reported for KITScenes Multimodal.

62 km2
HD map coverage

Production-grade Lanelet2 map area across three European cities.

1.78
UniDAC far-range depth δ1

δ1 score for monocular depth estimation beyond 200 m.

80%
Novel-view recall degradation

More than 80 percent relative traffic-sign recall loss at ±3 m lateral offset.

98.3%
Epona collision-free rate

Best reported open-loop collision-free rate at the 3-second horizon.

What the paper found

Researchers at Karlsruhe Institute of Technology and FZI Research Center for Information Technology introduce KITScenes Multimodal, a European autonomous-driving dataset designed to test spatial reasoning in difficult urban environments. Its synchronized robotaxi platform combines 72.5 Mpx global-shutter cameras, seven lidars reaching a measured maximum range of 409.2 m, three 4D imaging radars, and redundant GNSS/INS localization. Covering 62 km2 across Karlsruhe, Frankfurt, and Sindelfingen, the release pairs sensor data with reprojection-accurate Lanelet2 HD maps that encode road topology, regulatory elements, and 3D traffic lights, signs, and poles, and are validated in the open-source Autoware stack. Four benchmarks target complete online HD-map construction, long-range monocular depth, map-grounded novel-view synthesis, and multimodal end-to-end driving. The results expose failures hidden by conventional datasets: beyond 200 m, UniDAC achieves the strongest tested depth estimate but only a δ1 score of 1.78, while current novel-view synthesis loses more than 80 percent of traffic-sign recall at lateral offsets of ±3 m. In open-loop driving, the strongest Epona configuration reaches a 98.3 percent collision-free rate at the 3-second horizon. KITScenes is smaller than large corpora such as NVIDIA’s PhysicalAI AV, but prioritizes sensor fidelity, geographic diversity, complete maps, and deployment-oriented evaluation for Level 4 autonomy.

Original abstract

Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400m, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, validated through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving. Project page: https://kitscenes.com/

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis