NTH

SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

AuthorsPeiyu Liu, Dingxi Zhang, Federico Tombari, Marc Pollefeys, Christina Tsalicoglou, Daniel Barath

September 19, 2026 2 min read
Watch on YouTube
The one-line take

SplashSplat reconstructs fleeting real-world liquid splashes in 3D by combining multi-camera masks, coarse fluid motion, and dynamic Gaussian primitives.

Key results

20
Benchmark scenes

Real splash scenes captured for the synchronized multi-view benchmark.

60
Capture frame rate

Frames per second for the real-world liquid videos.

20.13
Real novel-view PSNR

SplashSplat's PSNR on held-out camera views.

37.8
Training time

Minutes required on a single NVIDIA V100 GPU, including preprocessing and reconstruction.

What the paper found

SplashSplat tackles the difficult problem of reconstructing opaque, weakly textured splashes whose sheets, ligaments, and droplets appear and disappear within milliseconds. The authors introduce the first real-world synchronized benchmark with 20 scenes captured by seven calibrated cameras at 4K resolution and 60 fps, providing refined liquid masks, container meshes, and fixed novel-view splits. The method fuses multi-view masks into per-frame signed distance fields, estimates a coarse velocity field through level-set transport with approximate incompressibility, and advects Lagrangian carriers in a forecast–correct–resample loop. Each carrier decodes local Gaussian primitives, combining observation-guided motion with differentiable Gaussian rendering; this avoids unconstrained deformation and avoids a full fluid solver requiring unobserved inflow, volume, and boundary conditions. Against Deformable-3DGS, SpacetimeGaussians, and 4D-Scaffold-GS, SplashSplat achieves 20.13 PSNR, 0.9570 SSIM, and 0.0806 LPIPS on real novel views, while reducing temporal jitter error to 0.0276. It also evaluates on NeuroFluid, supports sub-frame temporal interpolation and material style transfer without re-optimization, and runs on a single NVIDIA V100 in 37.8 minutes using 5.31 GiB of memory. The dataset and pipeline could support visual effects, fluid-simulation validation, and embodied-AI perception, while remaining limited by missing bubbles, foam, specular refraction, and fully interpretable momentum dynamics.

Original abstract

A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis