🛰️ Daily AI Frontier
‹ back to 2026-09-03

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Research Multimodal & Generative

Ranking

Overall 91
Content 100
Popularity 70

Observed public metrics from 1 member.

Representative image for SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Merged summary

TL;DR - SolarWM is an open foundation for training interactive, long-horizon video world models across heterogeneous datasets and model backbones. It matters because it standardizes data, adaptation, training, and inference while releasing the full pipeline, weights, and recipes for reproducible research.

  • Unifies 1.43 million clips from 10 datasets under a frame-aligned schema containing observations, camera geometry, captions, quality metadata, selection records, and provenance.
  • Adapts four models ranging from 5B to 33B parameters, based on Wan2.2, LTX-2.5, and MiniMax-H3, while retaining each backbone’s native representations and objectives.
  • Uses a three-stage recipe: bidirectional adaptation, teacher-forced autoregressive initialization, and distribution-matching distillation.
  • Produces causal models supporting real-time, minutes-to-hours interactive rollouts despite training on sequences only five seconds long.

Sources (1)

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

arXiv cs.CV Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang 2026-09-02 arXiv:2609.02886
Public signals Hugging Face upvotes 116
Providers: Hugging Face · Upvotes 116 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:24:38.003430 UTC

TL;DR - SolarWM is an open foundation for training interactive, long-horizon video world models across heterogeneous datasets and model backbones. It matters because it standardizes data, adaptation, training, and inference while releasing the full pipeline, weights, and recipes for reproducible research.

  • Unifies 1.43 million clips from 10 datasets under a frame-aligned schema containing observations, camera geometry, captions, quality metadata, selection records, and provenance.
  • Adapts four models ranging from 5B to 33B parameters, based on Wan2.2, LTX-2.5, and MiniMax-H3, while retaining each backbone’s native representations and objectives.
  • Uses a three-stage recipe: bidirectional adaptation, teacher-forced autoregressive initialization, and distribution-matching distillation.
  • Produces causal models supporting real-time, minutes-to-hours interactive rollouts despite training on sequences only five seconds long.
item →