🛰️ Daily AI Frontier
‹ back to 2026-08-25

ReWorld: An Interactive World Model with Long-Horizon Memory

Research Multimodal & Generative

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Representative image for ReWorld: An Interactive World Model with Long-Horizon Memory

Merged summary

TL;DR - ReWorld is an interactive video world model that combines real-time control with long-horizon visual memory under a fixed inference budget. It can stream 704Ă—1280 worlds while recalling and regenerating previously visited views during minute-long rollouts.

  • Mixed local/global attention heads and randomized routing balance short-term action following with full-history learning.
  • A bounded KV cache and pose-indexed landmark bank retrieve memories near the current camera pose without retaining full-history attention.
  • Metric-aligned multi-source training and palindrome trajectories teach consistent physical controls and revisitation memory.
  • LoRA-based distillation reduces generation to four sampling steps, while evaluations report leading control fidelity and video quality against six recent models.

Sources (1)

ReWorld: An Interactive World Model with Long-Horizon Memory

arXiv cs.AI Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu, Yihua Du, Wei Wang, Tianyi Gui, Lianghua Huang, Yingcong Chen 2026-08-24 arXiv:2608.23565
Public signals Hugging Face upvotes 24
Providers: Hugging Face · Upvotes 24 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-24 14:34:16.074749 UTC

TL;DR - ReWorld is an interactive video world model that combines real-time control with long-horizon visual memory under a fixed inference budget. It can stream 704Ă—1280 worlds while recalling and regenerating previously visited views during minute-long rollouts.

  • Mixed local/global attention heads and randomized routing balance short-term action following with full-history learning.
  • A bounded KV cache and pose-indexed landmark bank retrieve memories near the current camera pose without retaining full-history attention.
  • Metric-aligned multi-source training and palindrome trajectories teach consistent physical controls and revisitation memory.
  • LoRA-based distillation reduces generation to four sampling steps, while evaluations report leading control fidelity and video quality against six recent models.
item →