🛰️ Daily AI Frontier
‹ back to 2026-07-24

Unified Video Dense Prediction from Disjoint Data

arXiv cs.CV Multimodal & Generative Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee 2026-07-23

TL;DR - UniD is a unified video model that predicts eight dense scene properties using disjoint, task-specific datasets. It avoids costly pseudo-labeling by distilling expert supervision into a diffusion-pretrained shared backbone.

  • Predicts depth, normals, semantics, boundaries, human parts, albedo, shading, and materials.
  • Lightweight task projectors transfer knowledge from per-task experts without overlapping annotations.
  • Pretrained diffusion priors help bridge domain gaps and generalize to unseen scene-task combinations.
  • Achieves competitive specialist-level performance with improved temporal and cross-task consistency.

view merged work →