SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos
AuthorsPeiyu Liu, Dingxi Zhang, Federico Tombari, Marc Pollefeys, Christina Tsalicoglou, Daniel Barath
Resources
SplashSplat reconstructs fleeting real-world liquid splashes in 3D by combining multi-camera masks, coarse fluid motion, and dynamic Gaussian primitives.
Key results
Real splash scenes captured for the synchronized multi-view benchmark.
Frames per second for the real-world liquid videos.
SplashSplat's PSNR on held-out camera views.
Minutes required on a single NVIDIA V100 GPU, including preprocessing and reconstruction.
What the paper found
SplashSplat tackles the difficult problem of reconstructing opaque, weakly textured splashes whose sheets, ligaments, and droplets appear and disappear within milliseconds. The authors introduce the first real-world synchronized benchmark with 20 scenes captured by seven calibrated cameras at 4K resolution and 60 fps, providing refined liquid masks, container meshes, and fixed novel-view splits. The method fuses multi-view masks into per-frame signed distance fields, estimates a coarse velocity field through level-set transport with approximate incompressibility, and advects Lagrangian carriers in a forecast–correct–resample loop. Each carrier decodes local Gaussian primitives, combining observation-guided motion with differentiable Gaussian rendering; this avoids unconstrained deformation and avoids a full fluid solver requiring unobserved inflow, volume, and boundary conditions. Against Deformable-3DGS, SpacetimeGaussians, and 4D-Scaffold-GS, SplashSplat achieves 20.13 PSNR, 0.9570 SSIM, and 0.0806 LPIPS on real novel views, while reducing temporal jitter error to 0.0276. It also evaluates on NeuroFluid, supports sub-frame temporal interpolation and material style transfer without re-optimization, and runs on a single NVIDIA V100 in 37.8 minutes using 5.31 GiB of memory. The dataset and pipeline could support visual effects, fluid-simulation validation, and embodied-AI perception, while remaining limited by missing bubbles, foam, specular refraction, and fully interpretable momentum dynamics.
Original abstract
A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.