NTH

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

AuthorsManwen Liao, Xinyu Lian, Jian Mao, Kaixu Chen, Li Luo, Jinghao Yan, Wanshui Gan, Qiao Yu, Weitian Zhang, Chunhua Shen, Guang Chen, Bo Dai, Xudong Xu, Zhaoyang Lyu

August 24, 2026 2 min read
Watch on YouTube
The one-line take

MegaParts uses compact discrete geometry tokens and long-context autoregressive modeling to generate highly complex 3D objects with hundreds of coherent parts.

Key results

300
Maximum parts

Largest object complexity supported by MegaParts.

256K
Maximum sequence length

Long-context token capacity used for complex part-aware objects.

10M
Part-mesh training data

Nearly 10M part meshes used to train the VQ-VAE tokenizer.

440K
Generation assets

Approximately 440K part-aware assets used to train the generation model.

43.40
MegaParts FID

Text-conditioned FID on PartObjaverse-Tiny, lower than Cube’s 55.58.

0.77
High-part-count success rate

Valid-output rate for bounding-box-conditioned objects with 200 to 300 parts.

What the paper found

MegaParts presents a token-efficient autoregressive framework for part-aware 3D generation, addressing the memory and sequence-length bottleneck that limits diffusion and conventional autoregressive methods to relatively simple objects. Its causal vector-quantized VQ-VAE encodes each component with an adaptive-length prefix, selected through a rate–distortion objective that allocates more tokens to geometrically complex parts; alternating VQ-VAE optimization, optimal-transport quantization, and stochastic gradient shortcuts stabilize the discrete representation. A fine-tuned Qwen3-8B model then generates object-level and part-level bounding boxes followed by local shape tokens in one structured sequence, using context and tensor parallelism for long-context training. The system scales to 300 parts and 256K tokens, trained with nearly 10M part meshes and approximately 440K part-aware assets, while preserving explicit structure for editing and articulation. On PartObjaverse-Tiny, MegaParts records a text-conditioned FID of 43.40, outperforming Cube’s 55.58, TRELLIS-text’s 54.81, and ShapeLLM-Omni’s 70.99. In held-out bounding-box-conditioned tests, it maintains a 0.77 success rate for objects with 200 to 300 parts, whereas diffusion baselines often fail from memory exhaustion. The pipeline uses Roblox’s Cube as a tokenizer starting point, Kimi-K2.5 for geometry captions, and was profiled on an NVIDIA A800, positioning long-context language-model generation as a scalable alternative to diffusion for highly compositional 3D assets.

Original abstract

Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token length and memory. We introduce MegaParts, a scalable autoregressive 3D generation framework to address this challenge by combining structured sequence modeling with a token-efficient vector-quantized shape tokenizer. Our tokenizer learns discrete latent representations for part-level geometry by minimizing token usage subject to high-fidelity reconstruction, enabling adaptive-length tokenization based on geometric complexity. On top of this compact representation, we train a large language model to generate object bounding boxes, part bounding boxes, and part shape tokens within a unified structured sequence. Combined with efficient long-context training strategy, our token-efficient formulation scales to objects with up to 300 parts and sequence lengths up to 256k tokens. This substantially extends the scale of part-aware 3D generation while preserving compositional structure and enabling fine-grained part-level control. Our method achieves higher mesh quality than baseline autoregressive and diffusion models, showing that compressed discrete part tokens improve not only scalability but also the achievable fidelity of generated geometry. These results suggest that LLM native token-efficient autoregressive modeling is a compelling alternative to diffusion for large-scale part-aware 3D generation. The project page is available at https://expmaster.github.io/megaparts_webpage.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis