NTH

Reflection-aware Generative Novel View Synthesis

AuthorsGeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

September 11, 2026 2 min read
Watch on YouTube
The one-line take

Ref-GeNVS helps generative vision systems create realistic new viewpoints of mirror-containing scenes by explicitly using reflections as additional virtual camera views.

Key results

8
Synthetic scenes

Number of synthetic mirror scenes used for evaluation.

3
Mirror-NeRF scenes

Number of real scenes used from the Mirror-NeRF dataset.

0.129
Real sparse-view DreamSim

Ref-GeNVS with SEVA; lower is better, compared with 0.275 for SEVA.

0.951
Real sparse-view CLIP

Ref-GeNVS with SEVA; higher is better, compared with 0.920 for SEVA.

15.27
Synthetic sparse-view PSNR

Ref-GeNVS with SEVA on synthetic data.

397
Inference time

Approximate seconds per scene on an NVIDIA RTX 6000 Ada Generation GPU, versus 130 seconds for SEVA.

What the paper found

Ref-GeNVS is a training-free method for reflection-aware generative novel view synthesis from one or a few images containing mirrors. Instead of treating a mirror as an ordinary surface, it estimates the mirror plane with DAM, Depth Anything 3, and RANSAC, reflects the input camera poses using Householder reflection, and represents each mirror image as two complementary views. Its two-stage pipeline first masks the mirror and applies Mirror-gated attention so only mirror-region tokens from reflected virtual views condition a multi-view diffusion backbone such as SEVA or MVGenMaster; it then uses Reflection injection, decode–flip–encode latent alignment, and SDEdit-style boundary guidance to synthesize a mirror surface consistent with the generated scene. Evaluation covers 8 synthetic scenes and 3 real scenes from the Mirror-NeRF dataset, including single-image and three-image input settings. With SEVA on real sparse-view data, Ref-GeNVS reaches DreamSim 0.129 and CLIP similarity 0.951, compared with 0.275 and 0.920 for SEVA, while its synthetic-data PSNR reaches 15.27. The method also improves single-image results over FlexWorld, VistaDream, and SEVA, and supports downstream reconstruction with Pi3. The trade-off is computation: on an NVIDIA RTX 6000 Ada Generation GPU, Ref-GeNVS requires about 397 seconds per scene versus 130 seconds for SEVA, because of automated preprocessing and two diffusion-generation stages.

Original abstract

We propose Ref-GeNVS, a training-free, reflection-aware method for generative novel view synthesis (NVS) in mirror scenes. Existing multi-view diffusion models often fail to recognize the mirror in the scene and cannot exploit reflected content for scene generation. To fix this issue without additional training, our key idea is to treat a mirror image as two complementary views. From input images, we estimate the mirror plane and reflect camera poses to form virtual views. Based on this virtual view setup, we propose a two-stage generation method consisting of Mirror-gated attention and Reflection injection, which enables reflection-consistent NVS by explicitly leveraging reflection relationships in a multi-view diffusion model. Ref-GeNVS inherits the strong generalizability of the multi-view diffusion backbone, while it does not require finetuning. On synthetic and real scenes including mirrors, Ref-GeNVS outperforms recent generative NVS methods by generating reflection-consistent and contextually coherent novel views, revealing scene structure visible only through mirrors. Project page: https://kim-geonu.github.io/Ref-GeNVS/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis