Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Ranking
Overall
84
Content
90
Popularity
70
Observed public metrics from 1 member.
Merged summary
TL;DR - Alaya-EVOKE is an interactive video world model that combines bounded denoiser context with external scene memory for open-ended generation. Long-horizon teacher supervision helps its three-step student resist content drift while remaining responsive.
- Stores camera-indexed scene geometry externally and retrieves only view-relevant information.
- Uses sparse and linear attention to scale teacher memory and compute linearly with sequence length.
- Transfers 30-second, self-forced rollouts to a three-step student without classifier-free guidance.
- Achieves state-of-the-art WBench results and generates 1.5-second chunks in 2.11 seconds on one H200 at 384Ă—640.
Sources (1)
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Public signals
Hugging Face upvotes 166
TL;DR - Alaya-EVOKE is an interactive video world model that combines bounded denoiser context with external scene memory for open-ended generation. Long-horizon teacher supervision helps its three-step student resist content drift while remaining responsive.
- Stores camera-indexed scene geometry externally and retrieves only view-relevant information.
- Uses sparse and linear attention to scale teacher memory and compute linearly with sequence length.
- Transfers 30-second, self-forced rollouts to a three-step student without classifier-free guidance.
- Achieves state-of-the-art WBench results and generates 1.5-second chunks in 2.11 seconds on one H200 at 384Ă—640.