The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
AuthorsRichard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert, Nils Rack, Kaiwen Wang, Fabian Konstantinidis, Julian Truetsch, Carlos Fernandez, Annika Bätz, Kevin Rösch, Marlon Steiner, Willi Poh, Yinzhe Shen, Royden Wagner, Felix Hauser, Dominik Strutz, Jaime Villa, Gleb Stepanov, Holger Caesar, Ömer Şahin Taş, Frank Bieder, Jan-Hendrik Pauls, Christoph Stiller
Resources
KITScenes Multimodal is a high-fidelity autonomous driving dataset that pairs rich sensors and detailed maps with new benchmarks for mapping, depth, view synthesis, and end-to-end driving.
Key results
Combined synchronized global-shutter camera resolution.
Measured maximum range reported for KITScenes Multimodal.
Production-grade Lanelet2 map area across three European cities.
δ1 score for monocular depth estimation beyond 200 m.
More than 80 percent relative traffic-sign recall loss at ±3 m lateral offset.
Best reported open-loop collision-free rate at the 3-second horizon.
What the paper found
Researchers at Karlsruhe Institute of Technology and FZI Research Center for Information Technology introduce KITScenes Multimodal, a European autonomous-driving dataset designed to test spatial reasoning in difficult urban environments. Its synchronized robotaxi platform combines 72.5 Mpx global-shutter cameras, seven lidars reaching a measured maximum range of 409.2 m, three 4D imaging radars, and redundant GNSS/INS localization. Covering 62 km2 across Karlsruhe, Frankfurt, and Sindelfingen, the release pairs sensor data with reprojection-accurate Lanelet2 HD maps that encode road topology, regulatory elements, and 3D traffic lights, signs, and poles, and are validated in the open-source Autoware stack. Four benchmarks target complete online HD-map construction, long-range monocular depth, map-grounded novel-view synthesis, and multimodal end-to-end driving. The results expose failures hidden by conventional datasets: beyond 200 m, UniDAC achieves the strongest tested depth estimate but only a δ1 score of 1.78, while current novel-view synthesis loses more than 80 percent of traffic-sign recall at lateral offsets of ±3 m. In open-loop driving, the strongest Epona configuration reaches a 98.3 percent collision-free rate at the 3-second horizon. KITScenes is smaller than large corpora such as NVIDIA’s PhysicalAI AV, but prioritizes sensor fidelity, geographic diversity, complete maps, and deployment-oriented evaluation for Level 4 autonomy.
Original abstract
Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400m, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, validated through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving. Project page: https://kitscenes.com/
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.