NTH

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

AuthorsIgor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai

September 11, 2026 2 min read
Watch on YouTube
The one-line take

Marigold V2 turns powerful diffusion models into fast, sharp, and more generalizable predictors of depth and other dense visual properties.

Key results

74K
Training data

Compact HyperSim and vKITTI mixture used for Marigold V2 training

2.8
ETH3D AbsRel

Zero-shot affine-invariant depth error on ETH3D

99.2
ETH3D delta1

Percentage of ETH3D pixels within a 1.25 depth ratio

0.352
HyperSim SEE3

Soft Edge Error demonstrating boundary and fine-detail preservation

9.6
2048x2048 latency

Single-pass inference latency on one 32GB GPU

29.3GB
2048x2048 memory

Peak GPU memory for high-resolution inference

What the paper found

Marigold V2 repurposes Huawei’s Qwen-Image-Edit-2509, an open-source image-editing Diffusion Transformer, into a single-pass monocular depth estimator using a cost-efficient two-stage protocol. The model uses 4-bit quantization and rank-128 QLoRA adapters, then applies iREPA-depth, which aligns internal DiT representations with DINOv3 features extracted from ground-truth depth rather than RGB images. A second stage introduces SinkLoss, a Sinkhorn optimal-transport objective that matches depth values within local 5×5 blocks, tolerating noisy or ambiguous pixels while sharpening boundaries and preserving fur, foliage, hair, and thin structures. Trained on a compact 74K-image mixture of HyperSim and vKITTI, the system reaches an AbsRel of 2.8 and a δ1 score of 99.2 on ETH3D, with reported AbsRel improvements of 16–26% over the previous best on KITTI and ETH3D. It also achieves a HyperSim SEE3 edge error of 0.352, outperforming Pixel-Perfect Depth and InfiniDepth. Despite its large generative backbone, single-pass VAE inference remains feasible at 2048×2048, requiring 9.6 seconds and 29.3GB of GPU memory, and the same recipe transfers to depth completion, see-through depth, surface normals, and albedo estimation.

Original abstract

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis