NTH

Bi-FlowGS: Bridging Generative View Completion and Gaussian Geometry through Bidirectional Flow Co-Refinement

AuthorsYuetong Wang, Jinsheng Quan, Yi Yang, Yawei Luo

September 19, 2026 2 min read
Watch on YouTube
The one-line take

Bi-FlowGS lets generated videos and 3D Gaussian scenes iteratively correct each other so sparse-view reconstructions become both more visually complete and geometrically faithful.

Key results

800
Training scenes

Scenes sampled from DL3DV-10K for video-restoration training.

19.06
Mip-NeRF 360 PSNR

Bi-FlowGS PSNR with 9 input views.

48.01%
V2G geometry gain

Average relative geometry improvement in the 3-view Mip-NeRF 360 setting.

0.71 dB
Cross-benchmark PSNR gain

Average gain over the strongest baselines across Tanks and Temples, DL3DV-Benchmark, and CO3D.

What the paper found

Bi-FlowGS targets Geometry Cheating in sparse-view 3D Gaussian Splatting, where plausible novel-view images can hide incorrect Gaussian positions and depths. Its key idea is an optical-flow feedback loop: Video-to-Geometry Flow Distillation, or V2G, extracts reliable temporal correspondences from diffusion-restored videos using WAFT and aligns them with differentiable geometry-induced flow, directly supervising Gaussian geometry. The reverse Geometry-to-Video module, G2V, injects 3DGS-derived flow, depth cues, and DINOv2 reference features into a CogVideoX-5B-I2V restoration backbone, producing more temporally consistent pseudo-views and stronger teacher flow. Trained on 800 DL3DV-10K scenes with 49-frame clips, the system improves Mip-NeRF 360 performance to 19.06 PSNR with 9 input views, while V2G delivers a 48.01% average geometry gain in the 3-view setting and adds only 21.71–26.51 ms per optimization step. Across Tanks and Temples, DL3DV-Benchmark, and CO3D, Bi-FlowGS achieves an average PSNR gain of 0.71 dB over the strongest baselines across nine sparse-view settings. The experiments use NVIDIA RTX PRO 6000 Blackwell GPUs and show that bidirectional co-refinement improves both rendering fidelity and cross-view geometric consistency.

Original abstract

Sparse-view 3D scene reconstruction with 3D Gaussian Splatting (3DGS) is inherently underconstrained. Plausible renderings can also coexist with erroneous Gaussian geometry, as errors in positions or depths may be concealed by opacity, scale, and appearance; we term this failure mode Geometry Cheating. Existing regularization methods constrain geometry but remain limited to observed views, while video-diffusion-based methods complete unseen views yet mainly use them as RGB pseudo-supervision, underusing motion and temporal priors and lacking explicit geometry supervision. We present Bi-FlowGS, which uses optical flow to bridge generative view completion and Gaussian geometry regularization. Our plug-and-play Video-to-Geometry Flow Distillation (V2G) distills temporal correspondence priors from restored videos into Gaussian geometry to alleviate Geometry Cheating. Conversely, Geometry-to-Video Flow-Guided Restoration (G2V) uses the current 3DGS geometry to guide temporally consistent video restoration, providing more reliable generative supervision. Together, V2G and G2V form an implicit bidirectional co-refinement process, enabling restored videos and the optimized 3DGS scene to iteratively improve each other. Experiments demonstrate improved rendering quality and geometric consistency across wide-baseline and unbounded 360° benchmarks.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis