🛰️ Daily AI Frontier
‹ back to 2026-09-04

趋境科技与摩尔线程达成战略合作,高品质 AI Token 国产异构方案性价比超越国际先进算力

量子位 Efficiency & Systems 量子位的朋友们 2026-09-04
Representative image for 趋境科技与摩尔线程达成战略合作,高品质 AI Token 国产异构方案性价比超越国际先进算力

TL;DR - QJ Technology and Moore Threads announced a production-deployed heterogeneous LLM inference solution that separates prefill and decode workloads across different accelerators. They claim the domestic setup delivers production-grade performance with better per-token cost efficiency than advanced international compute alternatives under equivalent service requirements.

  • Moore Threads MTT S5000 cards handle prefill and KV-cache generation, while high-bandwidth GPUs perform decode; QJ Technology’s PD technology coordinates the heterogeneous resources.
  • Reported production metrics include over 50 tokens per second on average, a KV-cache hit rate above 90%, 99.9% stability, and low time to first token.
  • A prefill pool of four to five MTT S5000 servers reportedly offers better input-token price-performance than international alternatives under the project’s production standards.
  • The companies plan to package the system as “Token Pod” clusters combining MTT S5000 hardware, the MUSA software stack, QJ Technology’s inference system, and its ATaaS operations platform.

view merged work →