🛰️ Daily AI Frontier
‹ back to 2026-08-16

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Research Multimodal & Generative

Ranking

Overall 84
Content 90
Popularity 70

Observed public metrics from 1 member.

Merged summary

TL;DR - Alaya-EVOKE is an interactive video world model that combines bounded denoiser context with external scene memory for open-ended generation. Long-horizon teacher supervision helps its three-step student resist content drift while remaining responsive.

  • Stores camera-indexed scene geometry externally and retrieves only view-relevant information.
  • Uses sparse and linear attention to scale teacher memory and compute linearly with sequence length.
  • Transfers 30-second, self-forced rollouts to a three-step student without classifier-free guidance.
  • Achieves state-of-the-art WBench results and generates 1.5-second chunks in 2.11 seconds on one H200 at 384Ă—640.

Sources (1)

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

arXiv cs.CV Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao 2026-08-13 arXiv:2608.13546
Public signals Hugging Face upvotes 166
Providers: Hugging Face · Upvotes 166 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:55.312206 UTC

TL;DR - Alaya-EVOKE is an interactive video world model that combines bounded denoiser context with external scene memory for open-ended generation. Long-horizon teacher supervision helps its three-step student resist content drift while remaining responsive.

  • Stores camera-indexed scene geometry externally and retrieves only view-relevant information.
  • Uses sparse and linear attention to scale teacher memory and compute linearly with sequence length.
  • Transfers 30-second, self-forced rollouts to a three-step student without classifier-free guidance.
  • Achieves state-of-the-art WBench results and generates 1.5-second chunks in 2.11 seconds on one H200 at 384Ă—640.
item →