CubePart: An Open-Vocabulary Part-Controllable 3D Generator
AuthorsYiheng Zhu, Kangle Deng, Jean-Philippe Fauconnier, Inaki Navarro, Daiqing Li, Ava Pun, Yinan Zhang, Peiye Zhuang, Xiaoxia Sun, Maneesh Agrawala, Kiran Bhat, Tinghui Zhou
Resources
CubePart lets users generate 3D objects from text while specifying the parts they want, making AI-generated assets more useful for games, animation, and simulation.
Key results
Stage 1 is pretrained on approximately 4.7M mesh-text pairs.
The downscaled Qwen-VL-based Stage 1 model has 1.9B trainable parameters.
The open-vocabulary part-labeled training dataset contains 462K assets.
The same dataset contains 2.02M parts.
CubePart achieves 0.251 Chamfer Distance at the part level on PartObjaverse-Tiny.
CubePart achieves 0.974 F-score at the holistic level on PartObjaverse-Tiny.
What the paper found
CubePart, from Roblox and Stanford University collaborator Maneesh Agrawala, introduces an open-vocabulary, part-controllable 3D generator that turns a global text prompt plus a user-specified part schema into a coherent multi-mesh object, with each schema element becoming a distinct structurally complete part. The system is built on a two-stage rectified-flow diffusion architecture: Stage 1 trains a compact Qwen-VL-conditioned MM-DiT with 1.9B trainable parameters on about 4.7M mesh-text pairs to generate a holistic shape latent, then fine-tunes it with schema-aware prompts so requested parts are not omitted; Stage 2 reuses the Stage 1 weights and adds four zero-initialized cross-part attention blocks to decompose the full latent into part latents while preserving global coherence. To make this possible, the authors also construct a large open-vocabulary 3D part dataset of 462K assets and 2.02M parts, over 11× larger than PartVerse-XL, using 14-view Set-of-Mark rendering and GPT-5 to cluster and name parts automatically. On PartObjaverse-Tiny, CubePart outperforms prior part generators, reaching 0.251 Chamfer Distance and 0.743 F-score at the part level and 0.048 Chamfer Distance and 0.974 F-score at the holistic level, and the generated assets can be plugged directly into Roblox-style game behavior scripts for driving, flight, and weapon effects without manual post-processing.
Original abstract
Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part decompositions that cannot be aligned with application-specific requirements. We present CubePart, a generative framework for open-vocabulary, part-controllable 3D mesh generation that exposes part structure as an explicit inference-time control signal. Given a global text prompt and a user-defined parts schema expressed as an open-ended list of part names, our method generates a set of meshes - one per schema element - that assemble into a coherent object while respecting the specified semantic structure. To enable this capability, we introduce a scalable data pipeline to construct a large open-vocabulary, part-labeled 3D dataset, along with a two-stage generative architecture that separates global shape synthesis from part-level decoding. We demonstrate that the resulting assets can be directly integrated into game engines and driven by animation and behavior scripts without manual post-processing. Project Page: https://cubepart.github.io/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.