NTH

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

AuthorsEungjune Shim, Hansol Lee, Eunjung Ju

July 26, 2026 2 min read
Watch on YouTube
The one-line take

DiffGI makes thin-shell 3D generation more precise and efficient by combining differentiable geometry images with latent diffusion.

Key results

0.461
GarmageSet Chamfer Distance

Full DiffGI-VAE reconstruction score, reported at ×10^-3 scale.

0.961
GarmageSet Normal Consistency

Full TSDF plus normal-rendering-loss reconstruction score.

23K
Image-conditioned mesh vertices

Average vertices in DiffGI garment meshes.

1.21
RTX 4070 inference latency

Seconds for image-conditioned generation.

3.22
RTX 4070 peak VRAM

Gigabytes required for image-conditioned generation.

What the paper found

DiffGI, developed at CLO Virtual Fashion, targets a central weakness of volumetric 3D generation: watertight implicit fields often thicken or merge thin, open-boundary surfaces such as garments. The method replaces binary geometry-image masks with a continuous 2D Truncated Signed Distance Function, preserving subpixel boundary positions during downsampling, and introduces Differentiable Marching Squares, whose analytical interpolation lets 3D surface losses backpropagate into the 2D representation. A DiffGI-VAE combines position and TSDF channels with a differentiable nvdiffrast normal-rendering loss, compressing UV-friendly surfaces into a 32×32×4 latent space; initialization from Stable Diffusion 1.5 mainly accelerates convergence. On ABO, containing 7,900 furniture scans, and GarmageSet, containing 14,801 garments, the full model reaches a GarmageSet Chamfer Distance of 0.461 × 10^-3 and Normal Consistency of 0.961, outperforming occupancy-based variants and prior geometry-image systems. For image-conditioned garment generation, DiffGI produces meshes averaging 23K vertices, compared with 109K for TRELLIS and 526K for GarmageNet, while achieving 1.35 × 10^-2 Chamfer Distance and 0.48 F1. Its latent diffusion models use DiT, DINOv2-Large conditioning, and flow matching; image-conditioned inference takes 1.21 seconds with 3.22 GB VRAM on an RTX 4070, and 8.52 seconds on a MacBook M4 CPU. The remaining limitations are chart seams, rounding at extreme sharp edges, and the absence of joint texture or material generation.

Original abstract

Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thin-shell and non-manifold geometries such as garments. Geometry image-based approaches offer a surface-centric alternative, but existing methods rely on discrete binary occupancy maps whose resolution-dependent boundary encoding causes staircase artifacts and information loss upon downsampling, while surface reconstruction remains a non-differentiable post-processing step disconnected from the learning pipeline. To address this, we propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that seamlessly integrates surface representation and geometric optimization. DiffGI replaces binary maps with a continuous 2D Truncated Signed Distance Function (TSDF), which encodes boundary position at subpixel precision within a fixed grid resolution, eliminating resolution-dependent staircase artifacts even under aggressive downsampling. Building on this continuous field, we introduce a differentiable Marching Squares algorithm based on analytical linear interpolation, allowing gradients from 3D surface losses to propagate back to the 2D latent space. Leveraging this differentiable pipeline, we train a DiffGI-VAE augmented with a geometry-aware normal rendering loss to compress complex 3D surfaces into an ultra-compact 32X32 latent space, and instantiate a transformer-based latent diffusion model with a flow-matching objective on top of this space for conditional 3D generation. Extensive experiments on garment and object datasets demonstrate that our method achieves superior reconstruction fidelity and boundary precision compared to prior geometry-image and voxel-based approaches, while requiring significantly fewer computational resources.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis