VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion
AuthorsZador Pataki, Paul-Edouard Sarlin, Marc Pollefeys
VidMap turns difficult, uncalibrated videos into reliable 3D camera trajectories by combining the best ideas from SLAM, SfM, temporal matching, and monocular depth.
Key results
VidMap’s uncalibrated full-trajectory score on LaMAR
VidMap’s uncalibrated score on robot sequences
VidMap’s uncalibrated translation-recall AUC
VidMap’s translation-recall AUC on uncalibrated EuRoC
LaMAR 100-meter W-AUC after removing depth from global positioning
What the paper found
VidMap, from researchers at ETH Zurich, Google, and the Microsoft Spatial AI Lab including Marc Pollefeys, combines SLAM’s temporal reasoning with offline SfM’s global optimization for long, uncalibrated videos. It selects motion-aware keyframes, propagates tracks with RoMa dense matching, detects distant revisits using MegaLoc, and preserves whether each observation is sequential or a loop closure. That provenance enables robust optimization to trust sequential edges while downweighting potentially aliased loop closures through different losses. VidMap also integrates GeoCalib camera initialization and monocular metric-depth priors with per-image scale optimization inside GLOMAP-style global positioning and bundle adjustment, stabilizing pure rotation, forward motion, weak parallax, and scale drift. On uncalibrated LaMAR, VidMap achieves 88.5% full-trajectory W-AUC, compared with 76.6% for ViPE; on uncalibrated CroCoDL robot sequences, it reaches 80.3% full W-AUC. On ETH3D-SLAM, its uncalibrated 5-centimeter AUC is 51.4%, and on challenging grayscale, fast-motion EuRoC it reaches 99.0% at 10 meters. The ablations show why the components matter: removing depth from global positioning reduces LaMAR’s 100-meter W-AUC to 13.6%, versus 89.3% for the complete system. The approach remains offline and dependent on learned matching, calibration, and depth priors, but its central contribution is a non-causal, temporally aware reconstruction pipeline that prevents drift without sacrificing global consistency.
Original abstract
Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.