🛰️ Daily AI Frontier
‹ back to 2026-08-10

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

Research World Models

Ranking

Overall 69
Content 80
Popularity 43

Observed public metrics from 1 member.

Representative image for UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

Merged summary

TL;DR - UniJEPA is a single joint-embedding predictive architecture that merges image-level (photometric) and video-level (temporal) self-supervised world modeling into one shared latent space, removing the need for separate task-specific JEPA recipes.

  • Combines a next-embedding prediction loss with a Gaussian regularizer into one end-to-end objective, claimed to be provably anti-collapse without EMA, stop-gradient, or pre-trained encoders — and with a single loss hyperparameter.
  • The shared latent space is said to support "controllable abstraction": photometric prediction yields invariant structure, while temporal prediction yields equivariant dynamics.
  • After action-conditioned post-training on offline trajectories, it does zero-shot planning by treating goal features as prediction targets.
  • Reported to match or surpass task-specific JEPAs (I-JEPA, V-JEPA 2, DINO-WM, etc.) on image, video, and control benchmarks, planning up to tens of times faster than generative world models at comparable accuracy.

Sources (1)

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

arXiv cs.CV An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian 2026-08-07 arXiv:2608.07409
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-23 14:16:29.136318 UTC

TL;DR - UniJEPA is a single joint-embedding predictive architecture that merges image-level (photometric) and video-level (temporal) self-supervised world modeling into one shared latent space, removing the need for separate task-specific JEPA recipes.

  • Combines a next-embedding prediction loss with a Gaussian regularizer into one end-to-end objective, claimed to be provably anti-collapse without EMA, stop-gradient, or pre-trained encoders — and with a single loss hyperparameter.
  • The shared latent space is said to support "controllable abstraction": photometric prediction yields invariant structure, while temporal prediction yields equivariant dynamics.
  • After action-conditioned post-training on offline trajectories, it does zero-shot planning by treating goal features as prediction targets.
  • Reported to match or surpass task-specific JEPAs (I-JEPA, V-JEPA 2, DINO-WM, etc.) on image, video, and control benchmarks, planning up to tens of times faster than generative world models at comparable accuracy.
item →