🛰️ Daily AI Frontier
‹ back to 2026-08-05

114B参数、6B激活,Sand.ai刚刚开源全球首个千亿MoE视频生成模型

Industry & News Multimodal & Generative

Ranking

Overall 71
Content 80
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 114B参数、6B激活,Sand.ai刚刚开源全球首个千亿MoE视频生成模型

Merged summary

TL;DR - Sand.ai open-sourced MAGI-2-preview, a 114B-parameter MoE model for joint video and audio generation that activates about 6B parameters per forward pass. It provides researchers and enterprises with an unusually large open model for studying, fine-tuning, and privately deploying video MoE systems.

  • A single-stream Transformer jointly models text, video, and audio, enabling direct cross-modal interaction at every self-attention layer.
  • Multi-Head Latent MoE splits hidden states into 12 independently routed heads; each layer contains 3,072 expert units and activates 72 per token.
  • Sand.ai developed custom kernels and head-parallel execution to reduce routing overhead and keep communication independent of the number of activated experts.
  • The article reports sixth place on the AA video-generation leaderboard and an estimated cost of roughly ¥0.50 for a 10-second 1080p clip on eight H100 GPUs.

Sources (1)

114B参数、6B激活,Sand.ai刚刚开源全球首个千亿MoE视频生成模型

量子位 鹭羽 2026-08-05
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-04 14:20:29.226294 UTC

TL;DR - Sand.ai open-sourced MAGI-2-preview, a 114B-parameter MoE model for joint video and audio generation that activates about 6B parameters per forward pass. It provides researchers and enterprises with an unusually large open model for studying, fine-tuning, and privately deploying video MoE systems.

  • A single-stream Transformer jointly models text, video, and audio, enabling direct cross-modal interaction at every self-attention layer.
  • Multi-Head Latent MoE splits hidden states into 12 independently routed heads; each layer contains 3,072 expert units and activates 72 per token.
  • Sand.ai developed custom kernels and head-parallel execution to reduce routing overhead and keep communication independent of the number of activated experts.
  • The article reports sixth place on the AA video-generation leaderboard and an estimated cost of roughly ¥0.50 for a 10-second 1080p clip on eight H100 GPUs.
item →