🛰️ Daily AI Frontier
‹ back to 2026-08-24

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Research Human-Centric Vision

Ranking

Overall 75
Content 90
Popularity 39

Observed public metrics from 1 member.

Representative image for Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Merged summary

TL;DR - Human-JEPA is a video-pretrained vision model designed to both perceive people in the present and anticipate their future behavior. Its anchored forecasting approach preserves dense perception capabilities while adding anticipation, offering a unified alternative to separate static and predictive models.

  • Anchors dense prediction targets to a frozen copy of the model initialization to prevent degradation of dense perception during video pretraining.
  • Uses a past-to-future training split instead of block masking, avoiding reported drops of five points in action recognition and 17 points in person re-identification.
  • With frozen probes, outperforms pixel-anchored specialists on pose estimation and person re-identification despite having 2.7Ă— fewer parameters.
  • The released predictor head adds anticipation without degrading it, though the model remains weaker on high-resolution dense parsing.

Sources (1)

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

arXiv cs.CV Hui Wei, Licai Sun, Guoying Zhao 2026-08-21 arXiv:2608.21160
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-12 14:21:39.100699 UTC

TL;DR - Human-JEPA is a video-pretrained vision model designed to both perceive people in the present and anticipate their future behavior. Its anchored forecasting approach preserves dense perception capabilities while adding anticipation, offering a unified alternative to separate static and predictive models.

  • Anchors dense prediction targets to a frozen copy of the model initialization to prevent degradation of dense perception during video pretraining.
  • Uses a past-to-future training split instead of block masking, avoiding reported drops of five points in action recognition and 17 points in person re-identification.
  • With frozen probes, outperforms pixel-anchored specialists on pose estimation and person re-identification despite having 2.7Ă— fewer parameters.
  • The released predictor head adds anticipation without degrading it, though the model remains weaker on high-resolution dense parsing.
item →