Unified Video Dense Prediction from Disjoint Data
Merged summary
TL;DR - UniD is a unified video model that predicts eight dense scene properties using disjoint, task-specific datasets. It avoids costly pseudo-labeling by distilling expert supervision into a diffusion-pretrained shared backbone.
- Predicts depth, normals, semantics, boundaries, human parts, albedo, shading, and materials.
- Lightweight task projectors transfer knowledge from per-task experts without overlapping annotations.
- Pretrained diffusion priors help bridge domain gaps and generalize to unseen scene-task combinations.
- Achieves competitive specialist-level performance with improved temporal and cross-task consistency.
Sources (1)
Unified Video Dense Prediction from Disjoint Data
TL;DR - UniD is a unified video model that predicts eight dense scene properties using disjoint, task-specific datasets. It avoids costly pseudo-labeling by distilling expert supervision into a diffusion-pretrained shared backbone.
- Predicts depth, normals, semantics, boundaries, human parts, albedo, shading, and materials.
- Lightweight task projectors transfer knowledge from per-task experts without overlapping annotations.
- Pretrained diffusion priors help bridge domain gaps and generalize to unseen scene-task combinations.
- Achieves competitive specialist-level performance with improved temporal and cross-task consistency.