🛰️ Daily AI Frontier
‹ back to 2026-09-26

谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架

Industry & News Efficiency & Systems

Ranking

Overall 75
Content 85
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架

Merged summary

TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.

  • The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
  • Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
  • Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
  • The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.

Sources (1)

谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架

量子位 听雨 2026-09-26
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:13:29.561746 UTC

TL;DR - Inferact open-sourced a TPU-specific “megakernel” that reportedly runs Kimi K3 on 16 Google TPU v7 chips at 709 tokens/s, 57% faster than 16 Nvidia GB200 chips under the described benchmark. The result highlights how inference software and memory scheduling can outweigh headline hardware bandwidth.

  • The comparison used Kimi K3 and vLLM on both platforms; TPU v7 achieved 709 tokens/s versus GB200’s 452 tokens/s with DeepSeek’s DSpark speculative decoding.
  • Without speculative decoding, TPU throughput remained higher: 249 versus 127 tokens/s at batch size 1, and 865 versus 636 tokens/s at batch size 8.
  • Inferact fused Kimi K3’s 92 MoE layers into one Pallas program, reducing kernel-launch gaps and enabling cross-layer weight prefetching to keep memory bandwidth utilized.
  • The optimization preserved reported benchmark accuracy and cut compilation time from over 30 minutes with XLA to under 90 seconds, but is currently customized for Kimi K3’s architecture.
item →