NTH

CubePart: An Open-Vocabulary Part-Controllable 3D Generator

AuthorsYiheng Zhu, Kangle Deng, Jean-Philippe Fauconnier, Inaki Navarro, Daiqing Li, Ava Pun, Yinan Zhang, Peiye Zhuang, Xiaoxia Sun, Maneesh Agrawala, Kiran Bhat, Tinghui Zhou

June 10, 2026 2 min read
Watch on YouTube
The one-line take

CubePart lets users generate 3D objects from text while specifying the parts they want, making AI-generated assets more useful for games, animation, and simulation.

Key results

4.7M
single-mesh pretraining pairs

Stage 1 is pretrained on approximately 4.7M mesh-text pairs.

1.9B
trainable parameters

The downscaled Qwen-VL-based Stage 1 model has 1.9B trainable parameters.

462K
part dataset assets

The open-vocabulary part-labeled training dataset contains 462K assets.

2.02M
part dataset parts

The same dataset contains 2.02M parts.

0.251
PartObjaverse-Tiny part-level CD

CubePart achieves 0.251 Chamfer Distance at the part level on PartObjaverse-Tiny.

0.974
PartObjaverse-Tiny holistic F-score

CubePart achieves 0.974 F-score at the holistic level on PartObjaverse-Tiny.

What the paper found

CubePart, from Roblox and Stanford University collaborator Maneesh Agrawala, introduces an open-vocabulary, part-controllable 3D generator that turns a global text prompt plus a user-specified part schema into a coherent multi-mesh object, with each schema element becoming a distinct structurally complete part. The system is built on a two-stage rectified-flow diffusion architecture: Stage 1 trains a compact Qwen-VL-conditioned MM-DiT with 1.9B trainable parameters on about 4.7M mesh-text pairs to generate a holistic shape latent, then fine-tunes it with schema-aware prompts so requested parts are not omitted; Stage 2 reuses the Stage 1 weights and adds four zero-initialized cross-part attention blocks to decompose the full latent into part latents while preserving global coherence. To make this possible, the authors also construct a large open-vocabulary 3D part dataset of 462K assets and 2.02M parts, over 11× larger than PartVerse-XL, using 14-view Set-of-Mark rendering and GPT-5 to cluster and name parts automatically. On PartObjaverse-Tiny, CubePart outperforms prior part generators, reaching 0.251 Chamfer Distance and 0.743 F-score at the part level and 0.048 Chamfer Distance and 0.974 F-score at the holistic level, and the generated assets can be plugged directly into Roblox-style game behavior scripts for driving, flight, and weapon effects without manual post-processing.

Original abstract

Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part decompositions that cannot be aligned with application-specific requirements. We present CubePart, a generative framework for open-vocabulary, part-controllable 3D mesh generation that exposes part structure as an explicit inference-time control signal. Given a global text prompt and a user-defined parts schema expressed as an open-ended list of part names, our method generates a set of meshes - one per schema element - that assemble into a coherent object while respecting the specified semantic structure. To enable this capability, we introduce a scalable data pipeline to construct a large open-vocabulary, part-labeled 3D dataset, along with a two-stage generative architecture that separates global shape synthesis from part-level decoding. We demonstrate that the resulting assets can be directly integrated into game engines and driven by animation and behavior scripts without manual post-processing. Project Page: https://cubepart.github.io/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis