Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
AuthorsSeokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
Resources
Self-Geometry improves 3D vision foundation models at test time by using camera and multi-view geometry as self-supervision without requiring ground-truth labels.
Key results
Relative improvement over the frozen π3 model.
Relative improvement using predicted camera poses.
Relative gain in unposed geometry F1.
Fraction of ETH3D TTA iterations with conflicting primary-loss gradients.
Largest adapter overhead relative to a pretrained VFM.
Minutes for up to 40 input views on an NVIDIA RTX PRO 6000.
What the paper found
Self-Geometry is a ground-truth-free, plug-and-play test-time adaptation pipeline for pretrained 3D vision foundation models such as VGGT, π3, and Depth Anything 3. Instead of relying on model-generated features or external teachers, it uses 2D pixel correspondences extracted by LightGlue as pseudo-ground truth. Its Geometric Disentanglement Optimization combines a point-to-point Multi-View Consistency loss for pose and depth with a depth-independent Epipolar Consistency loss for camera pose, while Gradient Disentanglement projects conflicting gradients apart. Frame Angular-Neighbor sampling selects views using scene-scale-invariant SO(3) rotation distances, and LoRA adapters update only attention-block QKV parameters while the backbone remains frozen. Across six models and four benchmarks—7Scenes, ETH3D, ScanNet++, and HiRoom—the method improves performance over frozen models and competing approaches such as Free-Geometry and TCO. On π3, mean pose AUC@3 rises by 8.3% and unposed geometry F1 by 5.1%; on HiRoom, the gains for DA3-Base and DA3-Small reach 85.2% and 70.4%. Gradient conflict occurs in 42.4% of TTA iterations, motivating the proposed disentanglement. LoRA adds only 0.7% to 3.4% parameter overhead, and adaptation completes in 2 minutes for up to 40 input views on a single NVIDIA RTX PRO 6000.
Original abstract
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, $π^3$, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.