NTH

ZipSplat: Fewer Gaussians, Better Splats

AuthorsAlexander Veicht, Sunghwan Hong, Dániel Baráth, Marc Pollefeys

June 18, 2026 2 min read
Watch on YouTube
The one-line take

ZipSplat makes fast 3D scene reconstruction more flexible by clustering visual tokens into fewer, smarter Gaussians, cutting representation cost while improving quality.

Key results

2.1
PSNR gain vs YoNoSplat on DL3DV

ZipSplat exceeds the best pose-free baseline on DL3DV

1.2
PSNR gain vs YoNoSplat on RealEstate10K

ZipSplat exceeds the best pose-free baseline on RealEstate10K

6
Gaussian reduction vs pixel-aligned methods

ZipSplat uses about 6× fewer Gaussians than pixel-aligned baselines

24
Gaussian reduction vs DA3 on DL3DV

Token decoder uses 24× fewer Gaussians than the per-pixel DA3 baseline

5
PSNR after test-time token optimization

Token optimization adds about 5 dB PSNR in roughly 3 seconds on a single RTX 4090

401
Rendering speedup at 192 views

Scaled compression reaches 401 FPS at 192 views

What the paper found

ZipSplat, from ETH Zürich and Microsoft with Marc Pollefeys as an author, is a feed-forward 3D Gaussian Splatting method that breaks the usual one-Gaussian-per-pixel design by decoding compact scene tokens into unconstrained 3D Gaussians. A multi-view backbone, instantiated with DA3-Giant, produces visual tokens that are compressed by k-means into scene tokens, refined by cross- and self-attention, and decoded by a lightweight MLP into 32 Gaussians per token; a one-directional Chamfer geometric loss, depth supervision, coupled initialization, and a progressive view/compression schedule stabilize free 3D placement. On DL3DV and RealEstate10K, ZipSplat reaches state-of-the-art pose-free novel view synthesis while using about 6× fewer Gaussians than pixel-aligned baselines, including a 24× reduction versus DA3 on DL3DV. It improves over YoNoSplat by 2.1 dB PSNR on DL3DV and 1.2 dB on RealEstate10K, and with test-time token optimization it adds about 5 dB more PSNR in roughly 3 seconds on a single RTX 4090. The model also generalizes zero-shot to Mip-NeRF360 and ScanNet++, where its pose-free variant reaches 21.72 and 18.01 PSNR at 32 views, respectively, and a pose-conditioned version climbs to 23.49 PSNR on ScanNet++. Because compression is set at inference, one trained model traces a continuous quality-efficiency curve: at 192 views, scaled compression reduces storage from about 183 MB to 9.3 MB and boosts rendering speed from 40 FPS to 401 FPS.

Original abstract

Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel, tying the representation budget to camera resolution rather than scene complexity. A flat wall and a richly textured object thus produce equally many Gaussians despite very different geometric needs. We propose ZipSplat, a token-based feed-forward model that decouples Gaussian placement from the pixel grid. A multi-view backbone extracts dense visual tokens, and k-means clustering compresses them into a compact set of scene tokens. Cross- and self-attention refine these tokens, and a lightweight MLP decodes each into a group of Gaussians with unconstrained 3D positions. Because clustering is applied at inference, a single trained model spans the quality-efficiency curve without retraining. ZipSplat operates without ground-truth poses or intrinsics, yet sets a new state of the art on DL3DV and RealEstate10K with ${\sim}6{\times}$ fewer Gaussians than pixel-aligned methods, surpassing the best pose-free baseline by 2.1dB and 1.2dB PSNR, respectively. It further generalizes zero-shot to Mip-NeRF360 and ScanNet++, outperforming all comparable baselines. Our project page is at ${\href{https://veichta.com/zipsplat}{https://veichta.com/zipsplat}}$.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis