🛰️ Daily AI Frontier
‹ back to 2026-07-24

Unified Video Dense Prediction from Disjoint Data

Research Multimodal & Generative

Ranking

Overall 69
Content 80
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - UniD is a unified video model that predicts eight dense scene properties using disjoint, task-specific datasets. It avoids costly pseudo-labeling by distilling expert supervision into a diffusion-pretrained shared backbone.

  • Predicts depth, normals, semantics, boundaries, human parts, albedo, shading, and materials.
  • Lightweight task projectors transfer knowledge from per-task experts without overlapping annotations.
  • Pretrained diffusion priors help bridge domain gaps and generalize to unseen scene-task combinations.
  • Achieves competitive specialist-level performance with improved temporal and cross-task consistency.

Sources (1)

Unified Video Dense Prediction from Disjoint Data

arXiv cs.CV Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee 2026-07-23 arXiv:2607.21592
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-13 10:16:18.834100 UTC

TL;DR - UniD is a unified video model that predicts eight dense scene properties using disjoint, task-specific datasets. It avoids costly pseudo-labeling by distilling expert supervision into a diffusion-pretrained shared backbone.

  • Predicts depth, normals, semantics, boundaries, human parts, albedo, shading, and materials.
  • Lightweight task projectors transfer knowledge from per-task experts without overlapping annotations.
  • Pretrained diffusion priors help bridge domain gaps and generalize to unseen scene-task combinations.
  • Achieves competitive specialist-level performance with improved temporal and cross-task consistency.
item →