谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架
TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.
- The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
- Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
- Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
- The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.