NTH

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

AuthorsOrest Kupyn, Goutam Bhat, Philipp Henzler, Fabian Manhardt, Christian Rupprecht, Federico Tombari

June 26, 2026 3 min read
Watch on YouTube
The one-line take

FLAT turns a single image into an explorable 3D scene by directly decoding surface triangles from diffusion latents, aiming for better geometry than Gaussian-based methods.

Key results

21.45
RealEstate10K PSNR

FLAT triangle variant on RealEstate10K

0.853
RealEstate10K normal cosine

Mean geometric quality for FLAT triangles

0.587
2DGS normal cosine

Mean geometric quality baseline on the same benchmark

23.01
Refined RealEstate10K PSNR

FLAT after optional 250-step test-time refinement

21.23
Opaque mesh RealEstate10K PSNR

Game-engine-compatible mesh converted from triangles

0.5M
Opaque mesh vertices

Vertex count of the converted triangle mesh

What the paper found

FLAT, developed by Google Research with Oxford VGG and TU Munich affiliations, is a feedforward image-to-3D scene generator that decodes explicit triangle splats directly from frozen video-diffusion latents in a single pass, instead of regressing volumetric 3D Gaussians. The key novelty is a ray-centered triangle parameterization: each token predicts depth, a constrained Cholesky-style 2D shape transform, and residual local rotations, which prevents degenerate triangles and stabilizes training. FLAT also replaces the standard triangle-splatting window with a product-based soft coverage function that extends support beyond boundaries and improves gradient flow. Trained on RealEstate10K and DL3DV with real and synthetic videos, using Uni3C built on Wan-2.1 and supervised by photometric, LPIPS, depth, and normal losses, FLAT’s triangle variant reaches a mean normal cosine of 0.853 versus 0.587 for 2DGS, while preserving competitive appearance. On RealEstate10K, it reports 21.45 PSNR, 0.710 SSIM, and 0.245 LPIPS; with the optional 250-step refinement, it rises to 23.01 PSNR and 0.790 SSIM. For game-engine compatibility, FLAT converts its triangle soup into an opaque mesh: the resulting meshes use 0.5M vertices and achieve 21.23 PSNR on RealEstate10K, outperforming 3DGS mesh extraction by more than 7 dB. The paper also provides a systematic comparison of 3DGS, 2DGS, and triangles under identical latent-decoding conditions, showing that triangles trade a small amount of PSNR for substantially sharper, more physically grounded geometry.

Original abstract

Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric representations suitable for downstream use. Current video diffusion models offer high-quality generation and implicitly encode multi-view geometric structure in latent space. However, existing feedforward latent scene decoders typically output volumetric 3D Gaussians that lack a well-defined surface, limiting their use in simulation or standard graphics pipelines. This motivates decoding surface-aligned primitives that are not only renderable but also closer to explicit geometric assets. We ask whether compressed video diffusion latents can be mapped directly to explicit surface primitives in a single pass. To this end, we introduce FLAT and, for the first time, show that triangle splats can be decoded directly from video diffusion latents. Compared with decoding 3D Gaussians, predicting flat primitives is notoriously more challenging due to high sensitivity to primitive orientations, oftentimes leading to poor gradient flow. FLAT solves with two key ingredients: a ray-centered rotation parameterization for triangle regression and a novel product window function that improves gradient flow during differentiable triangle rendering. On standard benchmarks, FLAT achieves significantly better geometric accuracy while maintaining competitive visual quality compared to state-of-the-art feedforward baselines. We further show that a lightweight test-time refinement step converts the predicted triangle soup into a fully opaque, game-engine-ready representation that supports real-time rendering. By evaluating 3DGS, 2DGS, and triangle splatting variants under an identical training setup, we provide the first systematic analysis of representation tradeoffs in feedforward scene generation. The project page is available at https://flat-splat.github.io

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis