NTH

SAM3D-Part: Interactive Part Selection and Generation from 3D Objects

AuthorsJiahao Chang, Dong Du, Wanhu Sun, Yujian Zheng, Chuanyu Pan, Bowen Zhao, Chongjie Ye, Yuanming Hu, Xiaoguang Han

September 18, 2026 2 min read
Watch on YouTube
The one-line take

SAM3D-Part lets users ask for specific components of a 3D object and generates complete, correctly placed meshes rather than decomposing the entire object.

Key results

0.641
Multi-part BBox IoU

Best source-alignment score on the multi-part test set.

0.053
Multi-part Chamfer Distance

Lowest geometric distance on the multi-part test set.

0.849
Multi-part F1-0.1

Best relaxed-threshold geometry score on the multi-part test set.

3.51T
Fused cross-attention FLOPs

Stage 1 cost after token fusion, 15% below the SAM3D baseline.

187k
Verified training assets

Curated assets drawn from Objaverse, Objaverse-XL, ABO, 3D-FUTURE, and HSSD.

What the paper found

SAM3D-Part turns selective 3D editing into an interactive generation task: a user clicks a rendered view, SAM2 produces a part mask, and the system generates that component as a complete, reusable mesh rather than a partial surface. Its key innovation is combining 2D prompts—RGB, mask, and point-map tokens from DINOv2—with source-mesh geometry encoded by Hunyuan3D-2.1 ShapeVAE through pixel-aligned channel fusion. Built on the two-stage sparse generation design of TRELLIS, Stage 1 predicts occupancy and dense per-voxel XYZ correspondences, enabling closed-form scale-and-translation alignment without ICP; Stage 2 refines high-resolution geometry. A voxelized part cache records earlier extractions to reduce conflicts during sequential multi-part queries. On the multi-part test set, SAM3D-Part reaches 0.641 BBox IoU, 0.053 Chamfer Distance, and 0.849 F1-0.1, outperforming OmniPart and Meta’s SAM-3D-Objects. Token fusion also makes the added 3D conditioning cheaper: cross-attention falls to 3.51 T FLOPs, 15% below the SAM3D baseline. Training data comprises 187k verified assets curated from Objaverse, Objaverse-XL, ABO, 3D-FUTURE, and HSSD. The method remains weaker on thin lattice structures and heavily overlapping or nested parts, where its 64^3 sparse representation loses detail.

Original abstract

Part-level control is essential for modern 3D asset creation, where objects are frequently edited, reused, animated, or fabricated through their individual components. In many such workflows, users need only several specific components rather than a complete object decomposition. However, existing 3D generation methods produce all parts regardless of user intent, while promptable 3D segmentation methods typically output partial surfaces instead of reusable complete meshes. In addition, image-conditioned part generators further struggle to preserve hidden geometry and accurate placement without directly conditioning on the source mesh. To address these problems, we present SAM3D-Part, a prompt-driven framework for selective part generation from input 3D object meshes. Given a source mesh and a part prompt, SAM3D-Part first encodes the source geometry into compact mesh features and aligns them with the rendered image, selective mask, and point-map observations via pixel-wise channel fusion. The fused representation conditions a feed-forward generative model to produce only the queried component as a completed mesh. To place the generated part back into the source coordinate frame, SAM3D-Part predicts dense per-voxel correspondences and estimates the part transformation from distributed spatial evidence rather than a single global pose code. For sequential multi-part queries, previously generated parts are stored in a part cache and reused as contextual constraints, reducing conflicts among independently requested components. Extensive experiments and ablations demonstrate that SAM3D-Part can significantly improve source alignment, reduce conditioning cost, and enable consistent selective part generation, achieving state-of-the-art. Code and weights will be available at https://github.com/Jiahao620/sam3d-part.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis