Unified Video Dense Prediction from Disjoint Data
AuthorsYihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee
Resources
UniD uses diffusion-informed distillation to teach one video model eight scene-understanding tasks from separate datasets.
Key results
UniD score versus 0.289 for DINOv3-H
UniD precision at 70 percent recall on SAW
Improved from 37.8 through projector-only fine-tuning
Frames per second with a 16-frame memory bank
Estimated peak memory, reduced from 650GB for per-pixel distillation
What the paper found
In a paper from Yihong Sun and colleagues at Adobe Research and Cornell University, UniD tackles a central problem in scene understanding: dense labels are fragmented across incompatible image and video datasets. Built on Stable Diffusion v2, the model predicts eight properties—depth, surface normals, semantic segmentation, instance boundaries, human parts, albedo, shading, and materials—without co-annotated images or pseudo-labels. UniD first trains task specialists on their respective datasets, then distills their latent representations into one shared video U-Net through lightweight task projectors. Extended Self-Attention supplies historical-frame context, while temporal gradient matching transfers stability even to tasks without video supervision. On out-of-distribution tests, UniD reduces albedo WHDR to 0.207 versus 0.289 for the DINOv3-H baseline, and reaches 94.6 shading precision at 70 percent recall. Latent reconstruction is weaker for classification, but projector-only fine-tuning raises Cityscapes semantic mIoU from 37.8 to 58.4. The unified model runs at 1.56 FPS with a 16-frame memory bank and cuts estimated training memory from 650GB to 78GB compared with per-pixel distillation. Overall, the work shows that diffusion-model visual priors can bridge disjoint domains while improving out-of-distribution robustness, temporal coherence, and cross-task alignment, although segmentation still needs specialized fine-tuning.
Original abstract
Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labeling. Our key insight is that the strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gaps introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves competitive performance against per-task specialists and multi-task baselines, with strong generalization to out-of-distribution scenarios and enhanced temporal and cross-task consistency. Code and video results are available at https://unid-video.github.io/.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.