From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
AuthorsZepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
Resources
A physics-aware data pipeline and fast video diffusion model remove distracting glass reflections while introducing a benchmark for measuring progress.
Key results
Paired videos used for full-reference evaluation.
In-the-wild videos used for human perceptual assessment.
Peak signal-to-noise ratio achieved by S2R-Removal.
Structural similarity achieved with the complete two-stage objective.
Milliseconds per frame at 832 × 480 resolution.
Performance after enabling all Physics-Grounded Augmentation variants.
What the paper found
This paper presents a closed-loop system for removing reflections from video, addressing the shortage of paired reflected and clean footage with S2R-Synthesis. Instead of simple RGB blending, it composes reflection structure in line-art space and uses Physics-Grounded Augmentation to model six glass effects: roughness blur, thickness ghosting, reflectance variation, partial coverage, static reflections, and planar reflection. A Wan2.1 video diffusion renderer, trained from FLUX.2-Klein-9B pseudo-reflections filtered with Qwen3-VL-8B-Instruct, then produces temporally coherent reflected videos with exact clean counterparts. The resulting S2R-Removal model adapts Wan2.1 through LoRA-based reflection-aware latent training, residual-derived intensity supervision, and a second stage using reconstruction, SSIM, and LeReS depth-consistency losses. It removes reflections in one denoising step rather than iterative diffusion sampling. Evaluation uses the new S2R-Bench, including 60 paired videos in S2R-Ref and 50 real-world videos in S2R-Real. On S2R-Ref, the method reaches 28.84 dB PSNR and 0.8594 SSIM, while inference takes 87.09 ms/frame. An ablation with full Physics-Grounded Augmentation raises OpenRR-1k performance to 34.13 dB PSNR, showing that physically motivated synthesis improves both training diversity and dereflection quality.
Original abstract
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.