🛰️ Daily AI Frontier
‹ back to 2026-07-31

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Research Multimodal & Generative

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Representative image for EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Merged summary

TL;DR - EgoGenesis generates controllable egocentric robot-manipulation videos using geometry-aware memory and action conditioning. Its synthetic trajectories improve real-robot generalization, especially for dual-arm tasks.

  • OAPM anchors generation to the first-frame 3D scene while refreshing recent state during long autoregressive rollouts.
  • A3D-RoPE injects camera-aware end-effector motion into skeleton-to-video cross-attention for precise action control.
  • Adding 400 synthetic trajectories to 400 real ones raises out-of-distribution success from 77% to 84% for single-arm tasks.
  • Dual-arm success improves from 53% to 70% with the same augmentation strategy.

Sources (1)

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

arXiv cs.CV Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He, Kaibo Yin, Xiaobing Tu, Yinggui Wang, Jinkui Ren, Xiantao Zhang, Shijian Wang, Jinghong Liu, Linfeng Zhang 2026-07-30 arXiv:2607.28243
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-30 14:28:54.474298 UTC

TL;DR - EgoGenesis generates controllable egocentric robot-manipulation videos using geometry-aware memory and action conditioning. Its synthetic trajectories improve real-robot generalization, especially for dual-arm tasks.

  • OAPM anchors generation to the first-frame 3D scene while refreshing recent state during long autoregressive rollouts.
  • A3D-RoPE injects camera-aware end-effector motion into skeleton-to-video cross-attention for precise action control.
  • Adding 400 synthetic trajectories to 400 real ones raises out-of-distribution success from 77% to 84% for single-arm tasks.
  • Dual-arm success improves from 53% to 70% with the same augmentation strategy.
item →