NTH

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

AuthorsSeokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh

August 17, 2026 2 min read
Watch on YouTube
The one-line take

Self-Geometry improves 3D vision foundation models at test time by using camera and multi-view geometry as self-supervision without requiring ground-truth labels.

Key results

8.3%
π3 mean pose AUC@3 improvement

Relative improvement over the frozen π3 model.

5.1%
π3 mean unposed geometry F1 improvement

Relative improvement using predicted camera poses.

85.2%
DA3-Base HiRoom geometry improvement

Relative gain in unposed geometry F1.

42.4%
Gradient conflict frequency

Fraction of ETH3D TTA iterations with conflicting primary-loss gradients.

3.4%
Maximum LoRA parameter overhead

Largest adapter overhead relative to a pretrained VFM.

2
Per-scene adaptation time

Minutes for up to 40 input views on an NVIDIA RTX PRO 6000.

What the paper found

Self-Geometry is a ground-truth-free, plug-and-play test-time adaptation pipeline for pretrained 3D vision foundation models such as VGGT, π3, and Depth Anything 3. Instead of relying on model-generated features or external teachers, it uses 2D pixel correspondences extracted by LightGlue as pseudo-ground truth. Its Geometric Disentanglement Optimization combines a point-to-point Multi-View Consistency loss for pose and depth with a depth-independent Epipolar Consistency loss for camera pose, while Gradient Disentanglement projects conflicting gradients apart. Frame Angular-Neighbor sampling selects views using scene-scale-invariant SO(3) rotation distances, and LoRA adapters update only attention-block QKV parameters while the backbone remains frozen. Across six models and four benchmarks—7Scenes, ETH3D, ScanNet++, and HiRoom—the method improves performance over frozen models and competing approaches such as Free-Geometry and TCO. On π3, mean pose AUC@3 rises by 8.3% and unposed geometry F1 by 5.1%; on HiRoom, the gains for DA3-Base and DA3-Small reach 85.2% and 70.4%. Gradient conflict occurs in 42.4% of TTA iterations, motivating the proposed disentanglement. LoRA adds only 0.7% to 3.4% parameter overhead, and adaptation completes in 2 minutes for up to 40 input views on a single NVIDIA RTX PRO 6000.

Original abstract

Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, $π^3$, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis