NTH

CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge

AuthorsKendong Liu, Yuxin Yao, Junhui Hou

August 17, 2026 2 min read
Watch on YouTube
The one-line take

CANIS uses generative 3D priors and semantic image cues to align arbitrarily oriented objects into a meaningful canonical pose.

Key results

85.67%
Point-MAE canonicalized mIoU

Instance mIoU on ShapeNetPart after CANIS canonicalization, versus 31.43% for rotated inputs.

22 s
CANIS runtime

Full per-object runtime on an NVIDIA RTX A6000.

What the paper found

CANIS addresses category-agnostic 3D canonicalization by converting an arbitrarily oriented point cloud into a shared upright and front-facing frame without canonicalization-specific training or category templates. Its two-stage pipeline uses SCPS, or shape-guided canonical proxy synthesis, and PCAR, or patch-to-cluster semantic anchor registration. First, CANIS renders four candidate views, selects the most informative one, and feeds it to the frozen orientation-aligned generator TRELLIS-OA, an extension of TRELLIS. A weak structural latent from the input preserves instance geometry while the generator supplies a canonical orientation. During synthesis, voxel-to-image cross-attention is recorded; the selected RGB-D view then acts as an image-semantic bridge, linking image patches to proxy voxel clusters and back-projecting them onto the original point cloud. PCAR uses these semantic anchors to gate FCGF correspondence search, reject semantically incorrect matches on symmetric parts, and estimate the final rigid rotation with SVD and maximal-clique consensus. On Toys4K’s 24 categories, CANIS reaches instance consistency IC 0.052 and category consistency CC 0.083, outperforming CaCa at 0.064 and 0.098. On 12 Objaverse-OA categories, it achieves IC 0.032 and CC 0.086. The method also raises Point-MAE part-segmentation instance mIoU from 31.43% on rotated inputs to 85.67% after canonicalization, while improving dense correspondence. On an NVIDIA RTX A6000, the full pipeline takes 22 s per object, demonstrating a practical tradeoff for generation-assisted 3D alignment.

Original abstract

Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D canonicalization ultimately requires a semantically meaningful orientation. To address this gap, we propose CANIS, a category-agnostic, generation-assisted framework that introduces the semantic orientation prior of a frozen image-to-3D generative model into 3D canonicalization, without canonicalization-specific training or category-specific templates. Specifically, CANIS first renders the input object from candidate viewpoints, selects an informative view, and generates a proxy in a canonical orientation. During generation, a sparse structural latent encoded from the input guides the proxy to preserve the geometry of an object. CANIS then uses the selected image as a semantic bridge between the input and the proxy. Image patches identify semantic regions on the proxy, and depth back-projection locates the corresponding regions on the input. The resulting semantic anchors constrain geometric matching, from which we estimate the rigid transformation that canonicalizes the input. Experiments on synthetic benchmarks validate CANIS and its key components, while qualitative results on partial observations and OmniObject3D suggest its applicability to incomplete and real-world scans. CANIS also improves downstream 3D classification, part segmentation, and dense correspondence under arbitrary rotations. Project page: https://kenkenzaii.github.io/Canis.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis