NTH

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

AuthorsSaad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez

July 19, 2026 2 min read
Watch on YouTube
The one-line take

SUFLECA teaches visual features to understand 3D geometry, enabling faster and more reliable CAD model alignment from a single image.

Key results

674K
Training images

NOC-supervised feature learning uses 674K images from 12 real and synthetic datasets.

42.3%
ScanNet25k instance accuracy

Instance-averaged alignment accuracy under the ScanNet25k NMS protocol.

0.53s
Per-object runtime

CAD alignment runtime per object instance.

2178MB
Peak GPU memory

Peak GPU memory used by SUFLECA during alignment.

What the paper found

SUFLECA, from researchers at the University of Luxembourg and Universidad de Zaragoza, addresses zero-shot CAD-to-image alignment: estimating an object’s 9D rotation, translation, and anisotropic scale from one RGB image. Its central idea is to convert appearance-oriented foundation features into geometry-aware descriptors by training a lightweight NOC head over 674K images from 12 real and synthetic datasets, using a frozen DUNE-B encoder and Dense Prediction Transformer. At inference, the NOC head is discarded, while compact 384-dimensional features support mutual k-nearest-neighbor matching, anisotropic-scale estimation with iteratively reweighted least squares, geometric consensus filtering, and RANSAC-based Procrustes registration. On ScanNet25k, SUFLECA reaches 33.4% category accuracy and 42.3% instance accuracy, exceeding ZeroCAD by 10.3 and 12.2 percentage points and even surpassing fully supervised baselines. On the DiffCAD split, it achieves 36.1% and 44.8% category and instance accuracy, more than doubling ZeroCAD’s performance. The system avoids iterative pose refinement and runs in 0.53 seconds per object with 2,178 MB of GPU memory on an NVIDIA RTX 5090. Experiments on occluded, inexact CAD retrieval using CO3D show stronger generalization than DINOv3-L, DUNE-B, and Diorama, although CAD retrieval remains the main bottleneck for a fully zero-shot pipeline.

Original abstract

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a weakly-supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on 674K images spanning 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable one-to-one CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 33.4%/42.3% category/instance accuracy, outperforming, with a smaller computational footprint, the strongest zero-shot baseline by 10.3/12.2 percentage points and, for the first time on this benchmark, even surpassing fully supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis