UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
Ranking
Overall
69
Content
80
Popularity
43
Observed public metrics from 1 member.
Merged summary
TL;DR - UniJEPA is a single joint-embedding predictive architecture that merges image-level (photometric) and video-level (temporal) self-supervised world modeling into one shared latent space, removing the need for separate task-specific JEPA recipes.
- Combines a next-embedding prediction loss with a Gaussian regularizer into one end-to-end objective, claimed to be provably anti-collapse without EMA, stop-gradient, or pre-trained encoders — and with a single loss hyperparameter.
- The shared latent space is said to support "controllable abstraction": photometric prediction yields invariant structure, while temporal prediction yields equivariant dynamics.
- After action-conditioned post-training on offline trajectories, it does zero-shot planning by treating goal features as prediction targets.
- Reported to match or surpass task-specific JEPAs (I-JEPA, V-JEPA 2, DINO-WM, etc.) on image, video, and control benchmarks, planning up to tens of times faster than generative world models at comparable accuracy.
Sources (1)
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - UniJEPA is a single joint-embedding predictive architecture that merges image-level (photometric) and video-level (temporal) self-supervised world modeling into one shared latent space, removing the need for separate task-specific JEPA recipes.
- Combines a next-embedding prediction loss with a Gaussian regularizer into one end-to-end objective, claimed to be provably anti-collapse without EMA, stop-gradient, or pre-trained encoders — and with a single loss hyperparameter.
- The shared latent space is said to support "controllable abstraction": photometric prediction yields invariant structure, while temporal prediction yields equivariant dynamics.
- After action-conditioned post-training on offline trajectories, it does zero-shot planning by treating goal features as prediction targets.
- Reported to match or surpass task-specific JEPAs (I-JEPA, V-JEPA 2, DINO-WM, etc.) on image, video, and control benchmarks, planning up to tens of times faster than generative world models at comparable accuracy.