NTH

Unified Video Dense Prediction from Disjoint Data

AuthorsYihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee

July 26, 2026 2 min read
Watch on YouTube
The one-line take

UniD uses diffusion-informed distillation to teach one video model eight scene-understanding tasks from separate datasets.

Key results

0.207
Albedo WHDR on IIW

UniD score versus 0.289 for DINOv3-H

94.6
Shading P@0.7

UniD precision at 70 percent recall on SAW

58.4
Cityscapes semantic mIoU after fine-tuning

Improved from 37.8 through projector-only fine-tuning

1.56
Streaming inference with memory

Frames per second with a 16-frame memory bank

78GB
Latent-distillation memory

Estimated peak memory, reduced from 650GB for per-pixel distillation

What the paper found

In a paper from Yihong Sun and colleagues at Adobe Research and Cornell University, UniD tackles a central problem in scene understanding: dense labels are fragmented across incompatible image and video datasets. Built on Stable Diffusion v2, the model predicts eight properties—depth, surface normals, semantic segmentation, instance boundaries, human parts, albedo, shading, and materials—without co-annotated images or pseudo-labels. UniD first trains task specialists on their respective datasets, then distills their latent representations into one shared video U-Net through lightweight task projectors. Extended Self-Attention supplies historical-frame context, while temporal gradient matching transfers stability even to tasks without video supervision. On out-of-distribution tests, UniD reduces albedo WHDR to 0.207 versus 0.289 for the DINOv3-H baseline, and reaches 94.6 shading precision at 70 percent recall. Latent reconstruction is weaker for classification, but projector-only fine-tuning raises Cityscapes semantic mIoU from 37.8 to 58.4. The unified model runs at 1.56 FPS with a 16-frame memory bank and cuts estimated training memory from 650GB to 78GB compared with per-pixel distillation. Overall, the work shows that diffusion-model visual priors can bridge disjoint domains while improving out-of-distribution robustness, temporal coherence, and cross-task alignment, although segmentation still needs specialized fine-tuning.

Original abstract

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labeling. Our key insight is that the strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gaps introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves competitive performance against per-task specialists and multi-task baselines, with strong generalization to out-of-distribution scenarios and enhanced temporal and cross-task consistency. Code and video results are available at https://unid-video.github.io/.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis