DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation
AuthorsEungjune Shim, Hansol Lee, Eunjung Ju
Resources
DiffGI makes thin-shell 3D generation more precise and efficient by combining differentiable geometry images with latent diffusion.
Key results
Full DiffGI-VAE reconstruction score, reported at ×10^-3 scale.
Full TSDF plus normal-rendering-loss reconstruction score.
Average vertices in DiffGI garment meshes.
Seconds for image-conditioned generation.
Gigabytes required for image-conditioned generation.
What the paper found
DiffGI, developed at CLO Virtual Fashion, targets a central weakness of volumetric 3D generation: watertight implicit fields often thicken or merge thin, open-boundary surfaces such as garments. The method replaces binary geometry-image masks with a continuous 2D Truncated Signed Distance Function, preserving subpixel boundary positions during downsampling, and introduces Differentiable Marching Squares, whose analytical interpolation lets 3D surface losses backpropagate into the 2D representation. A DiffGI-VAE combines position and TSDF channels with a differentiable nvdiffrast normal-rendering loss, compressing UV-friendly surfaces into a 32×32×4 latent space; initialization from Stable Diffusion 1.5 mainly accelerates convergence. On ABO, containing 7,900 furniture scans, and GarmageSet, containing 14,801 garments, the full model reaches a GarmageSet Chamfer Distance of 0.461 × 10^-3 and Normal Consistency of 0.961, outperforming occupancy-based variants and prior geometry-image systems. For image-conditioned garment generation, DiffGI produces meshes averaging 23K vertices, compared with 109K for TRELLIS and 526K for GarmageNet, while achieving 1.35 × 10^-2 Chamfer Distance and 0.48 F1. Its latent diffusion models use DiT, DINOv2-Large conditioning, and flow matching; image-conditioned inference takes 1.21 seconds with 3.22 GB VRAM on an RTX 4070, and 8.52 seconds on a MacBook M4 CPU. The remaining limitations are chart seams, rounding at extreme sharp edges, and the absence of joint texture or material generation.
Original abstract
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thin-shell and non-manifold geometries such as garments. Geometry image-based approaches offer a surface-centric alternative, but existing methods rely on discrete binary occupancy maps whose resolution-dependent boundary encoding causes staircase artifacts and information loss upon downsampling, while surface reconstruction remains a non-differentiable post-processing step disconnected from the learning pipeline. To address this, we propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that seamlessly integrates surface representation and geometric optimization. DiffGI replaces binary maps with a continuous 2D Truncated Signed Distance Function (TSDF), which encodes boundary position at subpixel precision within a fixed grid resolution, eliminating resolution-dependent staircase artifacts even under aggressive downsampling. Building on this continuous field, we introduce a differentiable Marching Squares algorithm based on analytical linear interpolation, allowing gradients from 3D surface losses to propagate back to the 2D latent space. Leveraging this differentiable pipeline, we train a DiffGI-VAE augmented with a geometry-aware normal rendering loss to compress complex 3D surfaces into an ultra-compact 32X32 latent space, and instantiate a transformer-based latent diffusion model with a flow-matching objective on top of this space for conditional 3D generation. Extensive experiments on garment and object datasets demonstrate that our method achieves superior reconstruction fidelity and boundary precision compared to prior geometry-image and voxel-based approaches, while requiring significantly fewer computational resources.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.