AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition
AuthorsShunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu
AdaptVPR uses controlled generative scene changes and verification to create harder same-place images, making visual place recognition more robust to weather, lighting, and occlusions.
Key results
Verified synthetic same-place hard positives in the AdaptCities training dataset.
Unique GSV-Cities reference images used to construct AdaptCities.
BoQ improvement on the SF-XL-Occlusion benchmark.
Acceptance rate achieved by the complete adaptive, verified generation procedure.
Average offline generation time in seconds per candidate with AdaptVPR.
What the paper found
AdaptVPR addresses a central weakness in visual place recognition: synthetic edits can look realistic while accidentally changing the place itself. Starting from Google Street View-derived GSV-Cities, it uses Qwen3-VL-4B-Instruct to parse scene attributes and a rule-based scheduler to select one of three routes: Global Appearance for weather, illumination, and time-of-day changes; Local Occlusion for vehicles or pedestrians; and Dual for combined shifts. Generated images are verified with SuperPoint and LightGlue feature matching, RANSAC-based geometric consistency, and CLIP appearance diversity, while Local and Dual failures receive up to three rounds of prompt refinement through Qwen-LightX2V. The resulting AdaptCities dataset contains 160K verified same-place hard positives generated from 88,989 reference images. AdaptVPR is model-agnostic: adding these samples to systems such as BoQ, SALAD, EDTformer, and DINOv2-based models leaves their architectures and inference unchanged, yet improves Recall@1 by up to 9.2%, with the largest gains under cross-season, nighttime, and occlusion shifts. For example, BoQ gains 9.2% on SF-XL-Occlusion and 8.3% on Nordland-star. Joint geometry-and-diversity filtering retains 65.6% of candidates and improves quality over unfiltered synthesis, while increasing average generation time to 46.5 seconds per candidate, an offline cost with no added retrieval latency.
Original abstract
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby place, yet its robustness is often degraded by domain shifts arising from illumination, weather, seasonal changes, and dynamic occlusions. One contributing factor is the limited appearance diversity of the same place in existing training data. To address this issue, we propose AdaptVPR, a route-aware generative augmentation framework that constructs same-place hard positives for robust VPR training. AdaptVPR first uses a vision language model to parse scene attributes and estimate editing feasibility, while a rule-based scheduler determines the generation route according to editability scores and risk constraints. The generation process is decomposed into three complementary routes: the Global Appearance Route introduces global scene changes in weather, illumination, and time of day; the Local Occlusion Route inserts plausible dynamic occluders; and the Dual Route combines both types of perturbations to produce more challenging appearance shifts. Each generated candidate is evaluated using a VPR-oriented verification scheme based on geometric consistency and appearance diversity, reducing the risk of structural drift while ensuring sufficient appearance variation. Global candidates are generated once and rejected if verification fails, while Local Occlusion and Dual candidates use verification feedback for limited prompt refinement and regeneration. Using this framework, we construct AdaptCities, containing 160K verified synthetic same-place hard positives. Experiments across multiple VPR baselines and vision foundation backbones show consistent gains on standard benchmarks and substantial improvements under challenging domain shifts, with R@1 gains of up to 9.2%. The source code and data resources are publicly available at https://github.com/chenshunpeng/AdaptVPR.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.