B300半年涨三倍;大厂集体绕开算力中转站;Infra公司多名核心高管离职;Token成本击穿圈人再变现丨算力情报局Vol.13
Ranking
Overall
54
Content
55
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Leifeng.com's compute-intelligence column reports on China's AI hardware market: NVIDIA B300 server prices tripled in six months amid hoarding and fraud, domestic inference chips are being disaggregated by transformer block, and the economics of token-based AI services are breaking the old internet growth playbook.
- B300 full-system spot prices in China rose from ~RMB 4M to over RMB 12M per unit since the Spring Festival, quoted "daily" rather than weekly; one East China firm pursuing 10,000+ units allegedly pressured suppliers into exclusivity, while some "in-stock" offers are scams to harvest buyers' procurement and deployment data.
- Domestic inference silicon is splitting the transformer: one AFN-architecture vendor has taped out an FFN/MoE chip (Attention chip still in design, to run heterogeneously with GPUs), while another is starting with a Prefill-only chip targeting above-H100 performance, LPDDR first and 3D stacking later.
- 3D-stacked chip startups are seeing valuations jump from hundreds of millions to billions of RMB in six months, but yield, stacking process, and supply chain still gate mass production; investors reportedly won't fund new inference architectures pre-tape-out and see room for only one or two training-side GPU winners.
- Large firms avoid third-party API relay services over data security, TPM/RPM concurrency limits, unauditable billing, and account-ban risk (ByteDance and Tencent were both burned); more broadly, per-user token cost makes "free tokens to acquire users, monetize later" economically unworkable.
Sources (1)
B300半年涨三倍;大厂集体绕开算力中转站;Infra公司多名核心高管离职;Token成本击穿圈人再变现丨算力情报局Vol.13
Public signals
N/A
TL;DR - Leifeng.com's compute-intelligence column reports on China's AI hardware market: NVIDIA B300 server prices tripled in six months amid hoarding and fraud, domestic inference chips are being disaggregated by transformer block, and the economics of token-based AI services are breaking the old internet growth playbook.
- B300 full-system spot prices in China rose from ~RMB 4M to over RMB 12M per unit since the Spring Festival, quoted "daily" rather than weekly; one East China firm pursuing 10,000+ units allegedly pressured suppliers into exclusivity, while some "in-stock" offers are scams to harvest buyers' procurement and deployment data.
- Domestic inference silicon is splitting the transformer: one AFN-architecture vendor has taped out an FFN/MoE chip (Attention chip still in design, to run heterogeneously with GPUs), while another is starting with a Prefill-only chip targeting above-H100 performance, LPDDR first and 3D stacking later.
- 3D-stacked chip startups are seeing valuations jump from hundreds of millions to billions of RMB in six months, but yield, stacking process, and supply chain still gate mass production; investors reportedly won't fund new inference architectures pre-tape-out and see room for only one or two training-side GPU winners.
- Large firms avoid third-party API relay services over data security, TPM/RPM concurrency limits, unauditable billing, and account-ban risk (ByteDance and Tencent were both burned); more broadly, per-user token cost makes "free tokens to acquire users, monetize later" economically unworkable.