Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
AuthorsJiangshan Wang, Zeqiang Lai, Jiayi Guo, Xin Yang, Xin Huang, Jiarui Chen, Ziheng Ouyang, Chunchao Guo, Xiangyu Yue
AffiliationsMMLab, CUHK · Tencent Hunyuan · Tsinghua University · Fudan University · Shanghai Innovation Institute · Nankai University
Tex-Zero shows that high-quality native 3D textures may be learned from cleverly structured 2D images instead of costly real 3D assets.
Key results
Approximate total images from SA-1B, BLIP3o-60k, and ShareGPT-4o.
Tex-Zero generation score; NaTex scores 0.0754.
Tex-Zero generation score; NaTex scores 27.74.
Tex-Zero generation score.
What the paper found
Tencent’s Tex-Zero asks whether native 3D texturing really needs textured 3D assets for training—and shows it can instead learn from images alone. It converts each image into a colored 3D plane, divides it into 16 patches, then rotates and aggregates them to create varied surface orientations and occlusions while preserving the image’s fine detail. A sparse 3D VAE and a FLUX-style multimodal diffusion transformer, trained with flow matching, share a latent space for image conditions and 3D textures. Training uses approximately 11.1 million images from SA-1B, BLIP3o-60k, and ShareGPT-4o; at inference, the system textures real geometry from multi-view references. On six-view generation, Tex-Zero achieves LPIPS 0.0340, PSNR 35.68, and SSIM 0.983, compared with NaTex at LPIPS 0.0754 and PSNR 27.74. The central result is that realistic object geometry is not essential supervision: image-derived samples can teach the model how 3D surfaces organize, occlude, and reveal appearance, while a unified image-and-texture representation helps preserve fine details.
Original abstract
Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.
Read the original paperMore in Generative Models
Browse all 66 papers →FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan
FrameMorrow helps long-video generators remember the right past frames by predicting what future content will need next.
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu
LIFT lets users guide not just how a video camera moves, but exactly what should appear and where in a future view.
RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.