趋境科技与摩尔线程达成战略合作,高品质 AI Token 国产异构方案性价比超越国际先进算力
TL;DR - QJ Technology and Moore Threads announced a production-deployed heterogeneous LLM inference solution that separates prefill and decode workloads across different accelerators. They claim the domestic setup delivers production-grade performance with better per-token cost efficiency than advanced international compute alternatives under equivalent service requirements.
- Moore Threads MTT S5000 cards handle prefill and KV-cache generation, while high-bandwidth GPUs perform decode; QJ Technology’s PD technology coordinates the heterogeneous resources.
- Reported production metrics include over 50 tokens per second on average, a KV-cache hit rate above 90%, 99.9% stability, and low time to first token.
- A prefill pool of four to five MTT S5000 servers reportedly offers better input-token price-performance than international alternatives under the project’s production standards.
- The companies plan to package the system as “Token Pod” clusters combining MTT S5000 hardware, the MUSA software stack, QJ Technology’s inference system, and its ATaaS operations platform.