NTH

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

AuthorsLeonhard Sommer, Artur Jesslen, Basavaraj Sunagad, Adam Kortylewski

June 10, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes 3D object understanding more fine-grained by teaching a model to infer consistent parts and shapes across object categories from a single image, while also releasing a new large benchmark to measure that skill.

Key results

178k
HouseCorr3D images

benchmark scale

280
HouseCorr3D instances

unique object instances

31.3
Morpheus 2D PCK@0.1

mean 2D correspondence accuracy on HouseCorr3D

43.7
Morpheus 3D modal PCK@0.1

mean 3D visible correspondence accuracy on HouseCorr3D

41.5
Morpheus 3D amodal PCK@0.1

mean 3D occlusion-aware correspondence accuracy on HouseCorr3D

What the paper found

This paper reframes category-level object understanding as monocular 3D correspondence in camera space, where the task is to predict the same semantic point across two RGB-D views without pose normalization. To make that measurable, the authors introduce HouseCorr3D, the first large-scale benchmark for this setting, built from 178k image pairs spanning 50 household categories and 280 object instances, with keypoints annotated directly on CAD meshes, plus amodal labels for occluded regions and explicit symmetry handling. They then propose Morpheus, a morphable-object prior that disentangles canonical shape, deformation, and 6D pose, so correspondences emerge implicitly from shared template topology rather than from explicit correspondence supervision. Morpheus uses a DINOv2 ViT-S encoder, a learned deformation field over a hybrid SDF-to-mesh template, and mesh transfer via barycentric coordinates in camera space. On HouseCorr3D, it reaches 31.3 PCK@0.1 for 2D evaluation, 43.7 for 3D modal correspondences, and 41.5 for 3D amodal correspondences, outperforming GenPose++ at 26.7, 37.0, and 34.3 respectively. On a filtered real-world subset of Omni6DPose, it generalizes to 44.7 PCK@0.1 in 2D and 34.8 in 3D, showing that semantic 3D alignment can be learned from geometric supervision alone, provided the model is anchored by a shared morphable prior.

Original abstract

Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estimation, current representations fail to capture the fine-grained semantics needed for reasoning about object parts, functions, and interactions. In this work, we study category-level 3D correspondence in camera space -- predicting, from a single image, 3D locations that remain consistent across instances within a category -- and show that it can emerge without explicit correspondence supervision by learning a shared morphable object prior. To enable research in this direction, we introduce HouseCorr3D, the first large-scale benchmark for monocular category-level 3D correspondence with 178k images across 50 household object categories, 280 unique instances, and 3D keypoint annotations directly on CAD models. Crucially, HouseCorr3D provides amodal correspondence labels for occluded regions and explicit symmetry annotations, addressing key limitations of existing datasets. We further propose Morpheus, a method that learns morphable category-level shape priors by disentangling canonical shape, deformation, and object pose. Through this shared canonical grounding, semantically meaningful 3D correspondences in camera space emerge implicitly. These emerging 3D correspondences set a new state of the art on HouseCorr3D, demonstrating that semantic 3D object understanding can arise without direct correspondence supervision. Data and code are publicly available at https://github.com/GenIntel/HouseCorr3D.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis