NTH

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

AuthorsChunshi Wang, Haohan Weng, Junliang Ye, Biwen Lei, Yang Li, Zibo Zhao, Zeqiang Lai, Kaiyi Zhang, Yunhan Yang, Zhuo Chen, Chunchao Guo, Yawei Luo

July 11, 2026 2 min read
Watch on YouTube
The one-line take

PolyFlow turns mesh generation into a parallel continuous process by embedding topology into a learnable state space, making artist-style 3D mesh synthesis faster and more controllable.

Key results

0.008
Toys4K CD

best Chamfer Distance on Toys4K

0.021
Toys4K HD

best Hausdorff Distance on Toys4K

43%
CD improvement vs BPT

reported improvement over the strongest autoregressive baseline

40%
HD improvement vs BPT

reported improvement over the strongest autoregressive baseline

0.9991
TopoEmbedder F1

edge reconstruction F1 at 32-dimensional topology embedding

5.88
4k vertex inference time

PolyFlow total inference time in seconds on a single NVIDIA A100 GPU

What the paper found

PolyFlow from Tencent Hunyuan reframes artist-style mesh generation as continuous flow matching by introducing a compact topology embedder that converts discrete mesh adjacency into a continuous per-vertex embedding recoverable through spacetime distance thresholding. Instead of autoregressively decoding tokens, the model denoises a unified vertex state [xyz, normals, topo] in parallel with a Flux-based DiT Transformer conditioned on frozen point-cloud features from Hunyuan3D-Omni, then reconstructs edges and faces exactly from the generated embeddings. This design removes sequential error accumulation, gives exact user control over output vertex count, and generates meshes in seconds rather than tens of seconds or minutes. On Toys4K, PolyFlow achieves the best fidelity among BPT, MeshAnythingV2, DeepMesh, and FastMesh, reaching 0.008 Chamfer Distance and 0.021 Hausdorff Distance, which the paper reports as 43% and 40% better than the strongest autoregressive baseline, BPT. The topology embedder is most effective at 32 dimensions, where edge reconstruction reaches 0.9991 F1, and inference scales to 4,000 vertices in 5.88 seconds on a single NVIDIA A100, compared with 36.72 seconds for FastMesh and 554.5 seconds for BPT.

Original abstract

Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induces substantial computational overhead, falling orders of magnitude slower than parallel generative models. On the other hand, while continuous diffusion and flow-matching methods support efficient parallel synthesis across a variety of domains, they cannot be directly applied to meshes: mesh connectivity is inherently discrete and incompatible with standard continuous noise injection and denoising operations. To resolve this fundamental incompatibility, we introduce a compact topology embedder that projects discrete mesh vertex positions and normals into continuous per-vertex embeddings, where the original discrete adjacency information can be faithfully recovered via spacetime distance thresholding. After pretraining and freezing this embedder, any raw mesh can be fully converted into a continuous per-vertex state space unifying position, normal, and implicit topological attributes. Built upon this novel continuous mesh representation, we present PolyFlow, a Transformer-based flow-matching framework that achieves fully parallel vertex state denoising conditioned on extracted point-cloud features. During inference, our model completes generation rapidly via an ODE solver, and supports explicit, precise control over output mesh resolution by directly specifying the target vertex count. Extensive evaluations on the Toys4K benchmark demonstrate that PolyFlow surpasses state-of-the-art autoregressive baselines in both Chamfer Distance and Hausdorff Distance.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis