商汤大装置支撑智谱 GLM-5.3-Flash上线,国产异构助力前沿智能进入普惠时代
TL;DR - Zhipu launched and open-sourced GLM-5.3-Flash, a 320B-parameter mixture-of-experts multimodal model, using SenseTime’s domestically powered heterogeneous inference infrastructure. The deployment suggests Chinese accelerator clusters can support frontier-model inference at large scale with competitive cost and efficiency.
- GLM-5.3-Flash activates 18B of its 320B parameters and scored 57 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 according to the article.
- Pre-release testing under the Ox-Alpha alias reportedly processed 62 trillion tokens on domestic chips through OpenCode and OpenRouter.
- SenseTime says system-level optimizations tripled end-to-end serving performance over the initial baseline, bringing hardware efficiency and per-token cost close to mainstream NVIDIA GPUs.
- Its heterogeneous inference approach assigns different inference stages to suitable chip architectures, reportedly delivering 1.25× the price-performance of NVIDIA H-series systems and 2.5× the token capacity of domestic homogeneous inference at equal cost.