🛰️ Daily AI Frontier
‹ back to 2026-09-10

Programmable World Model

Research Multimodal & Generative

Ranking

Overall 88
Content 95
Popularity 70

Observed public metrics from 1 member.

Representative image for Programmable World Model

Merged summary

TL;DR - Programmable World Model separates explicit world-state evolution from video generation, enabling persistent entities and programmable rules in long-running interactive environments. This architecture supports playable games with coherent mechanics while using a pretrained video model as the visual renderer.

  • An agent converts natural-language instructions into executable entity states and transition rules.
  • A lightweight engine maintains persistent global state, including off-screen entities and non-visual attributes.
  • State-augmented 3D oriented bounding boxes and camera trajectories are compiled into conditioning signals for video generation.
  • On CombatStateBench, the method achieves 94% Count Accuracy and 98% State Accuracy, outperforming existing interactive video world models.

Sources (1)

Programmable World Model

arXiv cs.CV Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang 2026-09-09 arXiv:2609.10540
Public signals Hugging Face upvotes 113
Providers: Hugging Face · Upvotes 113 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:21:38.635066 UTC

TL;DR - Programmable World Model separates explicit world-state evolution from video generation, enabling persistent entities and programmable rules in long-running interactive environments. This architecture supports playable games with coherent mechanics while using a pretrained video model as the visual renderer.

  • An agent converts natural-language instructions into executable entity states and transition rules.
  • A lightweight engine maintains persistent global state, including off-screen entities and non-visual attributes.
  • State-augmented 3D oriented bounding boxes and camera trajectories are compiled into conditioning signals for video generation.
  • On CombatStateBench, the method achieves 94% Count Accuracy and 98% State Accuracy, outperforming existing interactive video world models.
item →