NTH

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

AuthorsChin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Wei-Chen Chiu, Yu-Lun Liu

September 11, 2026 2 min read
Watch on YouTube
The one-line take

Scal3R keeps online 3D reconstruction stable over long videos by querying poses against multiple past keyframes instead of relying on a single distant reference.

Key results

1%
Trainable parameter fraction

Learnable pose tokens account for about 1% of the frozen backbone parameters.

8 hours
Fine-tuning time

Training converges on a single NVIDIA A100 GPU.

69.7
KITTI average ATE

Scal3R average absolute trajectory error, compared with 182.2 for TTT3R.

5.63
Virtual KITTI average ATE

Scal3R with the CUT3R backbone on Virtual KITTI.

12
Inference reference count

Number of relative pose references used at inference.

14.4 FPS
Runtime with loop closure

Full Scal3R throughput on KITTI with loop closure enabled.

What the paper found

Scal3R addresses the long-video collapse seen in online 3D reconstruction systems such as CUT3R and STream3R, which regress every camera pose against a distant first-frame anchor. It observes that local depth remains stable while global pose drifts, then replaces absolute regression with multi-reference relative pose querying. Lightweight learnable pose tokens, adding only about 1% of backbone parameters, are injected into a completely frozen backbone through asymmetric attention: pose tokens read image features, while image tokens remain unchanged for point-cloud generation. Each incoming frame queries multiple keyframes, and online Pose-Graph Optimization with iSAM2, keyframe selection, and DINOv2-SALAD loop closure integrates the relative constraints into a globally consistent trajectory. Trained on TartanAir using 4-view samples, Scal3R converges in 8 hours on a single NVIDIA A100 GPU and uses K=12 references at inference. On KITTI, it reaches an average ATE of 69.7 versus 182.2 for TTT3R, reducing error by over 60%; on Virtual KITTI, the CUT3R variant achieves 5.63 ATE. The method also reports state-of-the-art results across Sintel, TUM-Dynamic, ScanNet, and 7-Scenes, while preserving real-time operation at 14.4 FPS with loop closure. Its central contribution is shifting long-range reconstruction from fragile global extrapolation to locally reliable relative constraints aggregated by an online graph.

Original abstract

Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone's local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: https://linjohnss.github.io/scal3r/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis