Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Ranking
Overall
75
Content
90
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - Human-JEPA is a video-pretrained vision model designed to both perceive people in the present and anticipate their future behavior. Its anchored forecasting approach preserves dense perception capabilities while adding anticipation, offering a unified alternative to separate static and predictive models.
- Anchors dense prediction targets to a frozen copy of the model initialization to prevent degradation of dense perception during video pretraining.
- Uses a past-to-future training split instead of block masking, avoiding reported drops of five points in action recognition and 17 points in person re-identification.
- With frozen probes, outperforms pixel-anchored specialists on pose estimation and person re-identification despite having 2.7Ă— fewer parameters.
- The released predictor head adds anticipation without degrading it, though the model remains weaker on high-resolution dense parsing.
Sources (1)
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Human-JEPA is a video-pretrained vision model designed to both perceive people in the present and anticipate their future behavior. Its anchored forecasting approach preserves dense perception capabilities while adding anticipation, offering a unified alternative to separate static and predictive models.
- Anchors dense prediction targets to a frozen copy of the model initialization to prevent degradation of dense perception during video pretraining.
- Uses a past-to-future training split instead of block masking, avoiding reported drops of five points in action recognition and 17 points in person re-identification.
- With frozen probes, outperforms pixel-anchored specialists on pose estimation and person re-identification despite having 2.7Ă— fewer parameters.
- The released predictor head adds anticipation without degrading it, though the model remains weaker on high-resolution dense parsing.