Unlocking Dense Metric Depth Estimation in VLMs
AuthorsHanxun Yu, Xuan Qu, Yuxin Wang, Jianke Zhu, Lei ke
Resources
This work turns a vision-language model into a single-pass depth predictor that can output dense geometry and text together, aiming to make VLMs more useful for 3D understanding.
Key results
The 8B model achieves the best average δ1 across the 9 DepthVLM-Bench test sets for metric depth estimation.
The benchmark/training corpus aggregates indoor and outdoor depth data from eight public datasets.
A lightweight DPT-style depth head is attached to the pretrained VLM backbone and is reported as under 1% of the LLM.
All training images are normalized to this canonical focal length to reduce camera-induced metric ambiguity.
DepthVLM predicts a dense pixel-aligned depth map in one pass, much faster than Youtu-VL and DepthLM.
What the paper found
Unlocking Dense Metric Depth Estimation in VLMs introduces DepthVLM, a vision-language model that natively predicts full-resolution metric depth maps and text responses in one forward pass by attaching a 34M-parameter DPT-style depth head to a pretrained LLM backbone such as Qwen3-VL-4B/8B. The key novelty is architectural and training, not distillation: instead of routing geometry through external 3D teachers or per-pixel queries, DepthVLM decodes multi-scale features from three ViT layers plus the LLM’s final image-token states, then trains in two stages, first freezing the VLM to fit only the depth head with SILog loss, then end-to-end fine-tuning with joint language and depth supervision. The paper also normalizes all training images to a shared focal length, fc = 1000, to remove camera-induced metric ambiguity, and releases DepthVLM-Bench, a 4.4M-image indoor-outdoor benchmark spanning Argoverse2, Waymo, DDAD, NuScenes, ScanNet++, Taskonomy, HM3D, and Matterport3D. On the paper’s δ1 metric, DepthVLM-8B reaches 0.876 average across 9 test sets, versus 0.730 for DepthLM-12B and 0.603 for Youtu-VL-4B, while also surpassing specialized pure vision models, including DepthAnythingV3 at 0.877 average and Metric3Dv2 at 0.812, with markedly better efficiency: 0.42 seconds per 256×192 image compared with 2.48 seconds for Youtu-VL and about 13 hours for DepthLM’s per-pixel query protocol. Importantly, the model preserves general VLM ability on benchmarks such as MMBench, ScienceQA, OCRBench, and POPE, and its improved dense geometry even boosts downstream 3D spatial reasoning.
Original abstract
Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained visual perception and prevents the recovery of dense geometry. Prior methods either distill geometry from external vision models, introducing error accumulation, or enable direct prediction with inefficient per-pixel query or coarse token-level outputs. In this paper, we propose DepthVLM, a simple yet effective framework that transforms a single VLM into a native dense geometry predictor while preserving its multimodal capability. By attaching a lightweight depth head to the LLM backbone and training under a unified vision-text supervision paradigm with a two-stage schedule, DepthVLM generates full-resolution depth maps alongside language outputs in a single forward pass. We further introduce a unified indoor-outdoor metric depth benchmark in a VLM-compatible format. Experiments show that DepthVLM significantly outperforms existing VLMs with higher inference efficiency, surpasses leading pure vision models, and improves complex 3D spatial reasoning, moving toward a truly unified foundation model. All code and checkpoints will be publicly released.
Read the original paper