Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
AuthorsLeonhard Sommer, Artur Jesslen, Basavaraj Sunagad, Adam Kortylewski
This paper makes 3D object understanding more fine-grained by teaching a model to infer consistent parts and shapes across object categories from a single image, while also releasing a new large benchmark to measure that skill.
Key results
benchmark scale
unique object instances
mean 2D correspondence accuracy on HouseCorr3D
mean 3D visible correspondence accuracy on HouseCorr3D
mean 3D occlusion-aware correspondence accuracy on HouseCorr3D
What the paper found
This paper reframes category-level object understanding as monocular 3D correspondence in camera space, where the task is to predict the same semantic point across two RGB-D views without pose normalization. To make that measurable, the authors introduce HouseCorr3D, the first large-scale benchmark for this setting, built from 178k image pairs spanning 50 household categories and 280 object instances, with keypoints annotated directly on CAD meshes, plus amodal labels for occluded regions and explicit symmetry handling. They then propose Morpheus, a morphable-object prior that disentangles canonical shape, deformation, and 6D pose, so correspondences emerge implicitly from shared template topology rather than from explicit correspondence supervision. Morpheus uses a DINOv2 ViT-S encoder, a learned deformation field over a hybrid SDF-to-mesh template, and mesh transfer via barycentric coordinates in camera space. On HouseCorr3D, it reaches 31.3 PCK@0.1 for 2D evaluation, 43.7 for 3D modal correspondences, and 41.5 for 3D amodal correspondences, outperforming GenPose++ at 26.7, 37.0, and 34.3 respectively. On a filtered real-world subset of Omni6DPose, it generalizes to 44.7 PCK@0.1 in 2D and 34.8 in 3D, showing that semantic 3D alignment can be learned from geometric supervision alone, provided the model is anchored by a shared morphable prior.
Original abstract
Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estimation, current representations fail to capture the fine-grained semantics needed for reasoning about object parts, functions, and interactions. In this work, we study category-level 3D correspondence in camera space -- predicting, from a single image, 3D locations that remain consistent across instances within a category -- and show that it can emerge without explicit correspondence supervision by learning a shared morphable object prior. To enable research in this direction, we introduce HouseCorr3D, the first large-scale benchmark for monocular category-level 3D correspondence with 178k images across 50 household object categories, 280 unique instances, and 3D keypoint annotations directly on CAD models. Crucially, HouseCorr3D provides amodal correspondence labels for occluded regions and explicit symmetry annotations, addressing key limitations of existing datasets. We further propose Morpheus, a method that learns morphable category-level shape priors by disentangling canonical shape, deformation, and object pose. Through this shared canonical grounding, semantically meaningful 3D correspondences in camera space emerge implicitly. These emerging 3D correspondences set a new state of the art on HouseCorr3D, demonstrating that semantic 3D object understanding can arise without direct correspondence supervision. Data and code are publicly available at https://github.com/GenIntel/HouseCorr3D.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.