Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
AuthorsChin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Wei-Chen Chiu, Yu-Lun Liu
Resources
Scal3R keeps online 3D reconstruction stable over long videos by querying poses against multiple past keyframes instead of relying on a single distant reference.
Key results
Learnable pose tokens account for about 1% of the frozen backbone parameters.
Training converges on a single NVIDIA A100 GPU.
Scal3R average absolute trajectory error, compared with 182.2 for TTT3R.
Scal3R with the CUT3R backbone on Virtual KITTI.
Number of relative pose references used at inference.
Full Scal3R throughput on KITTI with loop closure enabled.
What the paper found
Scal3R addresses the long-video collapse seen in online 3D reconstruction systems such as CUT3R and STream3R, which regress every camera pose against a distant first-frame anchor. It observes that local depth remains stable while global pose drifts, then replaces absolute regression with multi-reference relative pose querying. Lightweight learnable pose tokens, adding only about 1% of backbone parameters, are injected into a completely frozen backbone through asymmetric attention: pose tokens read image features, while image tokens remain unchanged for point-cloud generation. Each incoming frame queries multiple keyframes, and online Pose-Graph Optimization with iSAM2, keyframe selection, and DINOv2-SALAD loop closure integrates the relative constraints into a globally consistent trajectory. Trained on TartanAir using 4-view samples, Scal3R converges in 8 hours on a single NVIDIA A100 GPU and uses K=12 references at inference. On KITTI, it reaches an average ATE of 69.7 versus 182.2 for TTT3R, reducing error by over 60%; on Virtual KITTI, the CUT3R variant achieves 5.63 ATE. The method also reports state-of-the-art results across Sintel, TUM-Dynamic, ScanNet, and 7-Scenes, while preserving real-time operation at 14.4 FPS with loop closure. Its central contribution is shifting long-range reconstruction from fragile global extrapolation to locally reliable relative constraints aggregated by an online graph.
Original abstract
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone's local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: https://linjohnss.github.io/scal3r/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.