NTH

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

AuthorsUlrich Prestel, Stefan Andreas Baumann, Nick Stracke, Björn Ommer

June 14, 2026 2 min read
Watch on YouTube
The one-line take

RayDer turns self-supervised novel view synthesis from brittle multi-module training into a single scalable transformer that learns from real-world video and performs strongly zero-shot on many benchmarks.

Key results

2.7M
SpatialVid size

training corpus used for scaling experiments

59M
smallest model

RayDer-XS parameter count

743M
largest model

RayDer-L parameter count

29.57
RE10K PSNR

best zero-shot RE10K novel view synthesis result reported for the full model

17.11
LLFF PSNR

open-set LLFF result for RayDer-L-5762

0.997
scaling fit R2

power-law fit quality for compute-data scaling

What the paper found

RayDer, from CompVis at LMU Munich and the Munich Center for Machine Learning, reframes self-supervised novel view synthesis as a single-model scaling problem by collapsing camera estimation, scene reconstruction, and rendering into one feed-forward transformer. The key move is a minimal dynamic state token that absorbs time-varying content during training, allowing the model to learn static-scene NVS from unconstrained real-world video without reconstructing dynamics at inference. Training on SpatialVid’s 2.7M videos, RayDer scales cleanly across four model sizes, from 59M to 743M parameters, and the compute-optimal frontier is fit extremely well by a single power law with R2 = 0.997. The final RayDer-L-5762 model reaches 17.11 PSNR on LLFF and 21.38 on the 3-view split, while beating E-RayZer on the open-set benchmark suite with strong zero-shot results across LLFF, DTU, CO3D, WildRGBD, Mip-NeRF 360, and Tanks & Temples. Importantly, replacing RayZer’s multi-network pipeline with one unified backbone improves both stability and quality, and random-order autoregression plus local high-resolution layers raise RE10K PSNR to 29.57 and pose-transfer accuracy to 0.92 R@10° and 0.90 T@30° on DL3DV-10K. The paper’s central claim is that abundant video becomes a scalable supervision source once dynamics are treated as nuisance variation rather than a reconstruction target.

Original abstract

Self-supervised novel view synthesis (NVS) remains challenging to scale, despite the abundance of video data, largely due to the brittleness of training on realistic videos and the hard-to-predict scaling behavior of multi-network system designs. We introduce RayDer, a unified, feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering into a single backbone, turning self-supervised NVS into a well-posed single-model scaling problem. A minimal dynamic state, treated as a nuisance factor, absorbs time-varying content and enables stable training on unconstrained real-world video. Importantly, RayDer keeps static-scene NVS as its target task: dynamic content is leveraged purely as scalable supervision, not reconstructed as in dynamic-scene (4D) NVS. Across multiple model sizes and orders of magnitude in data, RayDer exhibits clean power-law scaling with data and compute, and outperforms static-scene data mixtures. On a large number of benchmarks, RayDer achieves strong zero-shot open-set performance competitive with state-of-the-art supervised approaches. Project Page: https://compvis.github.io/rayder

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis