NTH

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

AuthorsZador Pataki, Paul-Edouard Sarlin, Marc Pollefeys

August 2, 2026 2 min read
Watch on YouTube
The one-line take

VidMap turns difficult, uncalibrated videos into reliable 3D camera trajectories by combining the best ideas from SLAM, SfM, temporal matching, and monocular depth.

Key results

88.5%
LaMAR full W-AUC

VidMap’s uncalibrated full-trajectory score on LaMAR

80.3%
CroCoDL robot full W-AUC

VidMap’s uncalibrated score on robot sequences

51.4
ETH3D-SLAM AUC at 5 cm

VidMap’s uncalibrated translation-recall AUC

99.0
EuRoC AUC at 10 m

VidMap’s translation-recall AUC on uncalibrated EuRoC

13.6%
No-depth global-positioning ablation

LaMAR 100-meter W-AUC after removing depth from global positioning

What the paper found

VidMap, from researchers at ETH Zurich, Google, and the Microsoft Spatial AI Lab including Marc Pollefeys, combines SLAM’s temporal reasoning with offline SfM’s global optimization for long, uncalibrated videos. It selects motion-aware keyframes, propagates tracks with RoMa dense matching, detects distant revisits using MegaLoc, and preserves whether each observation is sequential or a loop closure. That provenance enables robust optimization to trust sequential edges while downweighting potentially aliased loop closures through different losses. VidMap also integrates GeoCalib camera initialization and monocular metric-depth priors with per-image scale optimization inside GLOMAP-style global positioning and bundle adjustment, stabilizing pure rotation, forward motion, weak parallax, and scale drift. On uncalibrated LaMAR, VidMap achieves 88.5% full-trajectory W-AUC, compared with 76.6% for ViPE; on uncalibrated CroCoDL robot sequences, it reaches 80.3% full W-AUC. On ETH3D-SLAM, its uncalibrated 5-centimeter AUC is 51.4%, and on challenging grayscale, fast-motion EuRoC it reaches 99.0% at 10 meters. The ablations show why the components matter: removing depth from global positioning reduces LaMAR’s 100-meter W-AUC to 13.6%, versus 89.3% for the complete system. The approach remains offline and dependent on learned matching, calibration, and depth priors, but its central contribution is a non-causal, temporally aware reconstruction pipeline that prevents drift without sacrificing global consistency.

Original abstract

Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis