Every9D-21M: Large-Scale Real-World 9D Canonicalization of Everyday Objects
AuthorsLeonhard Sommer, Emil Akopyan, Adam Kortylewski
This work builds the largest real-world 9D pose dataset to date by mining millions of object-centric videos, enabling better training and evaluation of pose models for everyday objects.
Key results
Every9D-21M provides 9D pose annotations for 21.8M real-world images.
Fewer than 0.01% of images require manual 9D annotation.
The full dataset was produced with 199 hours of human effort.
Training on Every9D-21M improves coarse 30° accuracy on HANDAL by 20.7 percentage points over training on ImageNet3D.
What the paper found
Every9D-21M, from University of Freiburg and CISPA Helmholtz Center for Information Security, introduces a large-scale real-world benchmark for 9D object canonicalization and pose estimation built from 109,000 object-centric videos in uCO3D. The dataset contains 21.8 million annotated images across 700 everyday object categories, making it roughly two orders of magnitude larger than prior real-world 9D pose datasets in both image and category count. Its key technical contribution is a reference-based canonicalization pipeline: videos are clustered with DINOv3 embeddings, a medoid reference is manually annotated for only about 1,000 objects, and pose is propagated to the rest through cross-instance alignment that combines alpha-shape geometry, multi-view DINO features, RANSAC initialization, and gradient refinement with a geometry-appearance loss. Cross-category orientation rules encode semantic axes from intrinsic function and human-object interaction, enabling symmetry-aware evaluation. Human labor is sharply reduced to about 199 hours total, with fewer than 0.01% of images manually labeled and every propagated pose verified from multiple viewpoints. A baseline monocular model using LitePT and a DINO image encoder shows that training on Every9D-21M improves transfer to PASCAL3D+ and ImageNet3D, and generalizes much better to HANDAL than training on ImageNet3D alone; for example, the authors report a 16.6-point gain over OrientAnythingV2 on coarse 30° rotation accuracy in the symmetry-unaware setting, plus a 20.7-point gain on HANDAL.
Original abstract
Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervision. Most existing datasets either rely heavily on synthetic renderings or provide limited coverage of real-world objects: the largest real-world 9D pose dataset to date contains only 17K annotated objects across 9 categories. We address this gap with Every9D-21M, a dataset of 9D pose annotations for 21.8M real-world images from 109K object- centric videos spanning 700 everyday object categories - two orders of magnitude larger than prior real-world 9D pose benchmarks in both image and category count. To achieve this scale, we leverage object-centric videos by reconstructing object- level point clouds via multi-view geometry and aligning similar instances into a shared canonical coordinate frame. Canonical poses are manually annotated for only a small set of reference objects (fewer than 0.01% of all images) and propagated to the remaining instances via cross-instance alignment. All propagated canonical poses are then verified from multiple viewpoints. We further introduce cross-category orientation rules that induce category-level symmetries, enabling symmetry-aware evaluation. Beyond establishing dedicated training and evaluation splits as a benchmark for 9D pose foundation models, we show that training on Every9D-21M improves performance on ImageNet3D and PASCAL3D+, and generalizes to HANDAL substantially better than training on ImageNet3D. Data and code are available at https://github.com/GenIntel/Every9D.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.