🛰️ Daily AI Frontier
‹ back to 2026-09-20

一张3090就能跑!全栈国产模型,把AI办公搬到企业本地

Industry & News LLM Agents

Ranking

Overall 61
Content 65
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 一张3090就能跑!全栈国产模型,把AI办公搬到企业本地

Merged summary

TL;DR - China Telecom has open-sourced Xing4.0-29B-A4B, a lightweight MoE model designed for private enterprise coding and agent workflows. Its 4-bit version can run on a single RTX 3090, making local AI deployment practical where data cannot leave company infrastructure.

  • The model has 29B total parameters but activates about 4B per inference; 4-bit quantization reduces memory use to roughly 15GB.
  • It supports 256K context, extendable to 512K, and uses MLA, multi-token prediction, and mHC connections to improve long-context efficiency, generation speed, and training stability.
  • Training used Huawei Ascend 910C hardware and MindSpore/MindFormers, with reported system optimizations raising training throughput by about 96% over out-of-box performance.
  • It targets repository-level coding and long-running agent tasks, with integrations for agent frameworks and inference stacks including OpenCode, Claude Code, SGLang, and vLLM.

Sources (1)

一张3090就能跑!全栈国产模型,把AI办公搬到企业本地

量子位 henry 2026-09-20
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:01.172871 UTC

TL;DR - China Telecom has open-sourced Xing4.0-29B-A4B, a lightweight MoE model designed for private enterprise coding and agent workflows. Its 4-bit version can run on a single RTX 3090, making local AI deployment practical where data cannot leave company infrastructure.

  • The model has 29B total parameters but activates about 4B per inference; 4-bit quantization reduces memory use to roughly 15GB.
  • It supports 256K context, extendable to 512K, and uses MLA, multi-token prediction, and mHC connections to improve long-context efficiency, generation speed, and training stability.
  • Training used Huawei Ascend 910C hardware and MindSpore/MindFormers, with reported system optimizations raising training throughput by about 96% over out-of-box performance.
  • It targets repository-level coding and long-running agent tasks, with integrations for agent frameworks and inference stacks including OpenCode, Claude Code, SGLang, and vLLM.
item →