Vera: Identity-Faithful Human Subject-to-Video Generation
TL;DR - Vera is a human-centric subject-to-video framework designed to preserve identity across frames and correctly bind identities in multi-person scenes. It matters because existing methods can produce identity drift, attribute swapping, and excessive copying from reference images.
- Builds on a million-pair identity-aligned human image-video dataset created through person-level cross-clip retrieval.
- Uses Identity-Focal Masked Supervision to focus learning on identity-relevant regions while limiting irrelevant artifacts.
- Introduces Reference-Aware Layer-wise Attention in the DiT backbone to maintain stable identity cues across layers.
- Reported experiments show improved identity consistency, subject-role binding, and motion naturalness.