Bi-FlowGS: Bridging Generative View Completion and Gaussian Geometry through Bidirectional Flow Co-Refinement
AuthorsYuetong Wang, Jinsheng Quan, Yi Yang, Yawei Luo
Resources
Bi-FlowGS lets generated videos and 3D Gaussian scenes iteratively correct each other so sparse-view reconstructions become both more visually complete and geometrically faithful.
Key results
Scenes sampled from DL3DV-10K for video-restoration training.
Bi-FlowGS PSNR with 9 input views.
Average relative geometry improvement in the 3-view Mip-NeRF 360 setting.
Average gain over the strongest baselines across Tanks and Temples, DL3DV-Benchmark, and CO3D.
What the paper found
Bi-FlowGS targets Geometry Cheating in sparse-view 3D Gaussian Splatting, where plausible novel-view images can hide incorrect Gaussian positions and depths. Its key idea is an optical-flow feedback loop: Video-to-Geometry Flow Distillation, or V2G, extracts reliable temporal correspondences from diffusion-restored videos using WAFT and aligns them with differentiable geometry-induced flow, directly supervising Gaussian geometry. The reverse Geometry-to-Video module, G2V, injects 3DGS-derived flow, depth cues, and DINOv2 reference features into a CogVideoX-5B-I2V restoration backbone, producing more temporally consistent pseudo-views and stronger teacher flow. Trained on 800 DL3DV-10K scenes with 49-frame clips, the system improves Mip-NeRF 360 performance to 19.06 PSNR with 9 input views, while V2G delivers a 48.01% average geometry gain in the 3-view setting and adds only 21.71–26.51 ms per optimization step. Across Tanks and Temples, DL3DV-Benchmark, and CO3D, Bi-FlowGS achieves an average PSNR gain of 0.71 dB over the strongest baselines across nine sparse-view settings. The experiments use NVIDIA RTX PRO 6000 Blackwell GPUs and show that bidirectional co-refinement improves both rendering fidelity and cross-view geometric consistency.
Original abstract
Sparse-view 3D scene reconstruction with 3D Gaussian Splatting (3DGS) is inherently underconstrained. Plausible renderings can also coexist with erroneous Gaussian geometry, as errors in positions or depths may be concealed by opacity, scale, and appearance; we term this failure mode Geometry Cheating. Existing regularization methods constrain geometry but remain limited to observed views, while video-diffusion-based methods complete unseen views yet mainly use them as RGB pseudo-supervision, underusing motion and temporal priors and lacking explicit geometry supervision. We present Bi-FlowGS, which uses optical flow to bridge generative view completion and Gaussian geometry regularization. Our plug-and-play Video-to-Geometry Flow Distillation (V2G) distills temporal correspondence priors from restored videos into Gaussian geometry to alleviate Geometry Cheating. Conversely, Geometry-to-Video Flow-Guided Restoration (G2V) uses the current 3DGS geometry to guide temporally consistent video restoration, providing more reliable generative supervision. Together, V2G and G2V form an implicit bidirectional co-refinement process, enabling restored videos and the optimized 3DGS scene to iteratively improve each other. Experiments demonstrate improved rendering quality and geometric consistency across wide-baseline and unbounded 360° benchmarks.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.