NTH

MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation

AuthorsJiale Xu, Wang Zhao, Ying Shan

June 28, 2026 2 min read
Watch on YouTube
The one-line take

MeshWeaver makes 3D mesh generation more efficient and accurate by weaving vertices through sparse voxel geometry instead of relying on long, brittle token sequences.

Key results

800K
training corpus

meshes used to train the model

600M
model size

parameters in the LLaMA3-style transformer backbone

18%
compression ratio

state-of-the-art mesh tokenization efficiency

22%
baseline compression cap

prior coordinate-level tokenization methods

0.116
Chamfer Distance

best point-cloud-conditioned generation result on Toys4K

0.087
Hausdorff Distance

best point-cloud-conditioned generation result on Toys4K

What the paper found

MeshWeaver, from Tencent PCG’s ARC Lab, reframes autoregressive mesh generation as sparse-voxel-guided “surface weaving” and shifts prediction from next-coordinate to next-vertex, which shortens token streams and improves topology-aware reasoning. The core design combines a hierarchical sparse-voxel encoder with a multi-level vertex tokenization scheme: voxel features serve as vertex representations, cross-attention injects local geometry at each refinement level, and occupied voxels act as a generation scaffold that masks invalid predictions. Trained on an 800K-mesh corpus built from Objaverse++, ShapeNet, 3D-Future, HSSD, and ABO, the model uses a 24-layer LLaMA3-style transformer with 600M parameters and 7-bit quantization. On point-cloud-conditioned generation over the Toys4K benchmark, it reduces Chamfer Distance to 0.116 and Hausdorff Distance to 0.087, while reaching 0.732 Normal Consistency and 0.914 absolute Normal Consistency. Its vertex-level tokenization achieves a state-of-the-art 18% compression ratio, compared with prior methods capped at about 22%, and it scales to meshes with up to 16K faces. An ablation further shows that cross-attention KV caching increases inference throughput from 26.8 to 30.7 tokens/s, a 14.5% gain.

Original abstract

Autoregressive mesh generation has gained attention by tokenizing meshes into sequences and training models in a language-modeling fashion. However, existing approaches suffer from two fundamental limitations: (i) low tokenization efficiency, which yields long token sequences and prevents scaling to high-poly meshes, and (ii) absence of geometry-aware guidance, as generation is conditioned only on global shape embeddings rather than local surface cues. We introduce MeshWeaver, an autoregressive framework that treats mesh generation as a surface weaving process by directly predicting the next vertex instead of independent coordinates. At its core is a multi-level sparse-voxel encoder that injects geometric context into the generative process in three complementary ways: providing voxel features as vertex representations, guiding token prediction via cross-attention to voxel features, and serving as a structural scaffold that constrains generation around the input surface. Our hierarchical design enables coarse-to-fine vertex prediction in a single decoding step, while tightly coupling the generative model with 3D geometry. Extensive experiments demonstrate that MeshWeaver achieves a state-of-the-art compression ratio of 18%, can generate meshes with up to 16K faces, and significantly improves geometric fidelity over prior approaches.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis