SAM3D-Part: Interactive Part Selection and Generation from 3D Objects
AuthorsJiahao Chang, Dong Du, Wanhu Sun, Yujian Zheng, Chuanyu Pan, Bowen Zhao, Chongjie Ye, Yuanming Hu, Xiaoguang Han
SAM3D-Part lets users ask for specific components of a 3D object and generates complete, correctly placed meshes rather than decomposing the entire object.
Key results
Best source-alignment score on the multi-part test set.
Lowest geometric distance on the multi-part test set.
Best relaxed-threshold geometry score on the multi-part test set.
Stage 1 cost after token fusion, 15% below the SAM3D baseline.
Curated assets drawn from Objaverse, Objaverse-XL, ABO, 3D-FUTURE, and HSSD.
What the paper found
SAM3D-Part turns selective 3D editing into an interactive generation task: a user clicks a rendered view, SAM2 produces a part mask, and the system generates that component as a complete, reusable mesh rather than a partial surface. Its key innovation is combining 2D prompts—RGB, mask, and point-map tokens from DINOv2—with source-mesh geometry encoded by Hunyuan3D-2.1 ShapeVAE through pixel-aligned channel fusion. Built on the two-stage sparse generation design of TRELLIS, Stage 1 predicts occupancy and dense per-voxel XYZ correspondences, enabling closed-form scale-and-translation alignment without ICP; Stage 2 refines high-resolution geometry. A voxelized part cache records earlier extractions to reduce conflicts during sequential multi-part queries. On the multi-part test set, SAM3D-Part reaches 0.641 BBox IoU, 0.053 Chamfer Distance, and 0.849 F1-0.1, outperforming OmniPart and Meta’s SAM-3D-Objects. Token fusion also makes the added 3D conditioning cheaper: cross-attention falls to 3.51 T FLOPs, 15% below the SAM3D baseline. Training data comprises 187k verified assets curated from Objaverse, Objaverse-XL, ABO, 3D-FUTURE, and HSSD. The method remains weaker on thin lattice structures and heavily overlapping or nested parts, where its 64^3 sparse representation loses detail.
Original abstract
Part-level control is essential for modern 3D asset creation, where objects are frequently edited, reused, animated, or fabricated through their individual components. In many such workflows, users need only several specific components rather than a complete object decomposition. However, existing 3D generation methods produce all parts regardless of user intent, while promptable 3D segmentation methods typically output partial surfaces instead of reusable complete meshes. In addition, image-conditioned part generators further struggle to preserve hidden geometry and accurate placement without directly conditioning on the source mesh. To address these problems, we present SAM3D-Part, a prompt-driven framework for selective part generation from input 3D object meshes. Given a source mesh and a part prompt, SAM3D-Part first encodes the source geometry into compact mesh features and aligns them with the rendered image, selective mask, and point-map observations via pixel-wise channel fusion. The fused representation conditions a feed-forward generative model to produce only the queried component as a completed mesh. To place the generated part back into the source coordinate frame, SAM3D-Part predicts dense per-voxel correspondences and estimates the part transformation from distributed spatial evidence rather than a single global pose code. For sequential multi-part queries, previously generated parts are stored in a part cache and reused as contextual constraints, reducing conflicts among independently requested components. Extensive experiments and ablations demonstrate that SAM3D-Part can significantly improve source alignment, reduce conditioning cost, and enable consistent selective part generation, achieving state-of-the-art. Code and weights will be available at https://github.com/Jiahao620/sam3d-part.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.