Video Generative Models as Geometry Learner
AuthorsHaosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
Resources
GeoNeXt repurposes video generation models to learn depth and surface geometry jointly, achieving strong zero-shot results with far less training data.
Key results
Synthetic Hypersim and Virtual KITTI 2 RGB-depth-normal samples used for GeoNeXt training.
Zero-shot affine-invariant depth error achieved by GeoNeXt on KITTI.
Zero-shot depth accuracy achieved by GeoNeXt on KITTI.
Mean surface-normal angular error achieved on iBims-1.
Seconds for the 5-run ensemble at 768 by 768 resolution on an NVIDIA A5000.
What the paper found
GeoNeXt repurposes Stable Video Diffusion, built on Stable Diffusion’s latent VAE, as a unified learner for monocular depth and surface normals. Instead of treating geometry as an isolated image, it reformulates prediction as next-frame generation: an RGB image, depth map, and normal map are denoised together in lockstep, using the video model’s temporal attention and generative priors to preserve structural consistency and fine detail. Only the denoising U-Net is fine-tuned, while the VAE remains frozen, and training uses 59K synthetic RGB-depth-normal samples from Hypersim and Virtual KITTI 2. With just 5 denoising steps and a lightweight 5-run ensemble, GeoNeXt generalizes zero-shot to real datasets. On KITTI, it reaches an AbsRel of 8.2 and δ1 of 92.6, outperforming the unified image-diffusion model GeoWizard, which uses 208K training samples. For surface normals, GeoNeXt records a mean angular error of 16.4 on iBims-1 and 69.2 percent within 11.25 degrees. The method also demonstrates practical efficiency: on an NVIDIA A5000 at 768 by 768 resolution, the 5-run configuration takes 10.0 seconds and uses 1.5B parameters. Ablations show that jointly reconstructing the image, depth, and normals is important; removing image reconstruction raises KITTI-style depth errors, while separate depth-only and normal-only models perform worse. The resulting geometry supports controllable image generation, relighting, and mesh reconstruction, suggesting that video generators such as Stable Video Diffusion can serve not only as synthesis engines but also as compact 3D reasoning systems.
Original abstract
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.