谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架
Ranking
Overall
75
Content
85
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.
- The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
- Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
- Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
- The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.
Sources (1)
谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架
Public signals
N/A
TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.
- The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
- Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
- Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
- The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.