MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
AuthorsYufei Cai, Xuesong Niu, Hao Lu, Kun Gai, Kai Wu, Guosheng Lin
MetaView uses diffusion, implicit geometry, and metric depth to generate spatially consistent novel views from a single image—even under large camera movements.
Key results
MetaView’s Dense Matching Distance on the hardest DL3DV viewpoint-change split.
Overall Dense Matching Distance on the DL3DV test set.
Dense Matching Distance on the RealEstate10K test set.
Dense Matching Distance on the Sekai-Real-Walking-HQ test set.
Ablation result after removing the metric-scale z-axis from RoPE.
What the paper found
MetaView, from researchers at Nanyang Technological University, Kuaishou Technology’s Kolors Team, and The Hong Kong University of Science and Technology (Guangzhou), targets monocular novel-view synthesis from one image under large camera changes. Instead of reconstructing an explicit 3D scene or relying on unconstrained implicit generation, it combines DepthAnything3-Giant’s hierarchical geometry features with minimal metric-scale cues. Built on Qwen-Image-Edit’s MM-DiT diffusion transformer, MetaView adds parallel image–geometry attention and modifies RoPE to encode camera intrinsics, extrinsics, and a depth-based z-axis, reducing scale drift while preserving the backbone’s semantic priors. The authors also introduce Dense Matching Distance, or DMD, which evaluates dense cross-view correspondence and is more informative than pixel metrics during extrapolative viewpoint changes. On DL3DV, MetaView reaches a DMD of 20.74 on the hard split, compared with 41.07 for Gen3C; across complete test sets, it records 10.29 on DL3DV, 6.44 on RealEstate10K, and 6.69 on Sekai-Real-Walk-HQ. Ablations show that removing the z-axis increases DMD from 9.33 to 14.45, confirming that metric-scale encoding is central to camera control, while geometry tokens preserve fine-grained relative structure. The released implementation is associated with KlingAIResearch, and experiments use 40 sampling steps with classifier-free guidance.
Original abstract
Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.