🛰️ Daily AI Frontier
‹ back to 2026-08-27

神秘「牛来」模型果然是智谱!GLM首个原生多模态,还用的国产卡

Industry & News Multimodal & Generative

Ranking

Overall 64
Content 70
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 神秘「牛来」模型果然是智谱!GLM首个原生多模态,还用的国产卡

Merged summary

TL;DR - Zhipu released and open-sourced GLM-5.3 Flash, its first native multimodal GLM-5 model, previously tested anonymously as Ox Alpha. It targets affordable long-running agentic and visual tasks while serving production traffic on Chinese accelerator hardware.

  • The 320B-parameter mixture model activates 18B parameters across 45 layers and was pretrained on 30T multimodal tokens.
  • A hybrid linear-and-sparse attention architecture supports 1M-token contexts while reportedly cutting attention computation by 3.01× and KV-cache size by 4.44× versus GLM-5.3.
  • Demonstrations covered video understanding and editing, UI generation from design images, interactive 3D webpages, and a 12-hour autonomous Blender workflow.
  • Zhipu says its disaggregated encode-prefill-decode serving architecture tripled end-to-end performance on domestic chips; model weights, APIs, and coding integrations are now available.

Sources (1)

神秘「牛来」模型果然是智谱!GLM首个原生多模态,还用的国产卡

量子位 十三 2026-08-26
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:18:04.821269 UTC

TL;DR - Zhipu released and open-sourced GLM-5.3 Flash, its first native multimodal GLM-5 model, previously tested anonymously as Ox Alpha. It targets affordable long-running agentic and visual tasks while serving production traffic on Chinese accelerator hardware.

  • The 320B-parameter mixture model activates 18B parameters across 45 layers and was pretrained on 30T multimodal tokens.
  • A hybrid linear-and-sparse attention architecture supports 1M-token contexts while reportedly cutting attention computation by 3.01× and KV-cache size by 4.44× versus GLM-5.3.
  • Demonstrations covered video understanding and editing, UI generation from design images, interactive 3D webpages, and a 12-hour autonomous Blender workflow.
  • Zhipu says its disaggregated encode-prefill-decode serving architecture tripled end-to-end performance on domestic chips; model weights, APIs, and coding integrations are now available.
item →