🛰️ Daily AI Frontier
‹ back to 2026-09-26

谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架

量子位 Efficiency & Systems 听雨 2026-09-26
Representative image for 谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架

TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.

  • The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
  • Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
  • Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
  • The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.

view merged work →