SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
AuthorsSaad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez
SUFLECA teaches visual features to understand 3D geometry, enabling faster and more reliable CAD model alignment from a single image.
Key results
NOC-supervised feature learning uses 674K images from 12 real and synthetic datasets.
Instance-averaged alignment accuracy under the ScanNet25k NMS protocol.
CAD alignment runtime per object instance.
Peak GPU memory used by SUFLECA during alignment.
What the paper found
SUFLECA, from researchers at the University of Luxembourg and Universidad de Zaragoza, addresses zero-shot CAD-to-image alignment: estimating an object’s 9D rotation, translation, and anisotropic scale from one RGB image. Its central idea is to convert appearance-oriented foundation features into geometry-aware descriptors by training a lightweight NOC head over 674K images from 12 real and synthetic datasets, using a frozen DUNE-B encoder and Dense Prediction Transformer. At inference, the NOC head is discarded, while compact 384-dimensional features support mutual k-nearest-neighbor matching, anisotropic-scale estimation with iteratively reweighted least squares, geometric consensus filtering, and RANSAC-based Procrustes registration. On ScanNet25k, SUFLECA reaches 33.4% category accuracy and 42.3% instance accuracy, exceeding ZeroCAD by 10.3 and 12.2 percentage points and even surpassing fully supervised baselines. On the DiffCAD split, it achieves 36.1% and 44.8% category and instance accuracy, more than doubling ZeroCAD’s performance. The system avoids iterative pose refinement and runs in 0.53 seconds per object with 2,178 MB of GPU memory on an NVIDIA RTX 5090. Experiments on occluded, inexact CAD retrieval using CO3D show stronger generalization than DINOv3-L, DUNE-B, and Diorama, although CAD retrieval remains the main bottleneck for a fully zero-shot pipeline.
Original abstract
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a weakly-supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on 674K images spanning 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable one-to-one CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 33.4%/42.3% category/instance accuracy, outperforming, with a smaller computational footprint, the strongest zero-shot baseline by 10.3/12.2 percentage points and, for the first time on this benchmark, even surpassing fully supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.