🛰️ Daily AI Frontier
‹ back to 2026-09-19

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Research Multimodal & Generative

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Representative image for Astronex-World 1.0: Real-Time Interactive World Model Foundation

Merged summary

TL;DR - Astronex-World 1.0 is an open 5B-parameter video world model that supports controllable, persistent generation from text or images. It matters because it achieves real-time 832Ă—480 video at 24 fps on one NVIDIA L20 GPU while outperforming several larger models on WBench.

  • Supports frame-aligned camera trajectories, continuous actions, embodiment IDs, and text events inserted during rollouts.
  • Provides bidirectional full-context generation and block-causal generation with cross-block KV caching for persistent streaming.
  • Uses PRoPE for camera parameters and a 64-dimensional action stream that modulates every Transformer layer.
  • Its five-stage training pipeline runs on two 48 GB L20 GPUs; the final model scores 73.5 on WBench Navi and 70.0 on WBench Full.

Sources (1)

Astronex-World 1.0: Real-Time Interactive World Model Foundation

arXiv cs.CV Xin Zhou, Cong Miao 2026-09-17 arXiv:2609.20034
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:18:03.586970 UTC

TL;DR - Astronex-World 1.0 is an open 5B-parameter video world model that supports controllable, persistent generation from text or images. It matters because it achieves real-time 832Ă—480 video at 24 fps on one NVIDIA L20 GPU while outperforming several larger models on WBench.

  • Supports frame-aligned camera trajectories, continuous actions, embodiment IDs, and text events inserted during rollouts.
  • Provides bidirectional full-context generation and block-causal generation with cross-block KV caching for persistent streaming.
  • Uses PRoPE for camera parameters and a 64-dimensional action stream that modulates every Transformer layer.
  • Its five-stage training pipeline runs on two 48 GB L20 GPUs; the final model scores 73.5 on WBench Navi and 70.0 on WBench Full.
item →