RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
AuthorsUlrich Prestel, Stefan Andreas Baumann, Nick Stracke, Björn Ommer
Resources
RayDer turns self-supervised novel view synthesis from brittle multi-module training into a single scalable transformer that learns from real-world video and performs strongly zero-shot on many benchmarks.
Key results
training corpus used for scaling experiments
RayDer-XS parameter count
RayDer-L parameter count
best zero-shot RE10K novel view synthesis result reported for the full model
open-set LLFF result for RayDer-L-5762
power-law fit quality for compute-data scaling
What the paper found
RayDer, from CompVis at LMU Munich and the Munich Center for Machine Learning, reframes self-supervised novel view synthesis as a single-model scaling problem by collapsing camera estimation, scene reconstruction, and rendering into one feed-forward transformer. The key move is a minimal dynamic state token that absorbs time-varying content during training, allowing the model to learn static-scene NVS from unconstrained real-world video without reconstructing dynamics at inference. Training on SpatialVid’s 2.7M videos, RayDer scales cleanly across four model sizes, from 59M to 743M parameters, and the compute-optimal frontier is fit extremely well by a single power law with R2 = 0.997. The final RayDer-L-5762 model reaches 17.11 PSNR on LLFF and 21.38 on the 3-view split, while beating E-RayZer on the open-set benchmark suite with strong zero-shot results across LLFF, DTU, CO3D, WildRGBD, Mip-NeRF 360, and Tanks & Temples. Importantly, replacing RayZer’s multi-network pipeline with one unified backbone improves both stability and quality, and random-order autoregression plus local high-resolution layers raise RE10K PSNR to 29.57 and pose-transfer accuracy to 0.92 R@10° and 0.90 T@30° on DL3DV-10K. The paper’s central claim is that abundant video becomes a scalable supervision source once dynamics are treated as nuisance variation rather than a reconstruction target.
Original abstract
Self-supervised novel view synthesis (NVS) remains challenging to scale, despite the abundance of video data, largely due to the brittleness of training on realistic videos and the hard-to-predict scaling behavior of multi-network system designs. We introduce RayDer, a unified, feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering into a single backbone, turning self-supervised NVS into a well-posed single-model scaling problem. A minimal dynamic state, treated as a nuisance factor, absorbs time-varying content and enables stable training on unconstrained real-world video. Importantly, RayDer keeps static-scene NVS as its target task: dynamic content is leveraged purely as scalable supervision, not reconstructed as in dynamic-scene (4D) NVS. Across multiple model sizes and orders of magnitude in data, RayDer exhibits clean power-law scaling with data and compute, and outperforms static-scene data mixtures. On a large number of benchmarks, RayDer achieves strong zero-shot open-set performance competitive with state-of-the-art supervised approaches. Project Page: https://compvis.github.io/rayder
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.