Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
AuthorsIgor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai
Resources
Marigold V2 turns powerful diffusion models into fast, sharp, and more generalizable predictors of depth and other dense visual properties.
Key results
Compact HyperSim and vKITTI mixture used for Marigold V2 training
Zero-shot affine-invariant depth error on ETH3D
Percentage of ETH3D pixels within a 1.25 depth ratio
Soft Edge Error demonstrating boundary and fine-detail preservation
Single-pass inference latency on one 32GB GPU
Peak GPU memory for high-resolution inference
What the paper found
Marigold V2 repurposes Huawei’s Qwen-Image-Edit-2509, an open-source image-editing Diffusion Transformer, into a single-pass monocular depth estimator using a cost-efficient two-stage protocol. The model uses 4-bit quantization and rank-128 QLoRA adapters, then applies iREPA-depth, which aligns internal DiT representations with DINOv3 features extracted from ground-truth depth rather than RGB images. A second stage introduces SinkLoss, a Sinkhorn optimal-transport objective that matches depth values within local 5×5 blocks, tolerating noisy or ambiguous pixels while sharpening boundaries and preserving fur, foliage, hair, and thin structures. Trained on a compact 74K-image mixture of HyperSim and vKITTI, the system reaches an AbsRel of 2.8 and a δ1 score of 99.2 on ETH3D, with reported AbsRel improvements of 16–26% over the previous best on KITTI and ETH3D. It also achieves a HyperSim SEE3 edge error of 0.352, outperforming Pixel-Perfect Depth and InfiniDepth. Despite its large generative backbone, single-pass VAE inference remains feasible at 2048×2048, requiring 9.6 seconds and 29.3GB of GPU memory, and the same recipe transfers to depth completion, see-through depth, surface normals, and albedo estimation.
Original abstract
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.