蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-Flash
Merged summary
TL;DR - Ant Group’s Bailing team released Ling-3.0-Flash, a 124B-parameter hybrid-reasoning model that activates 5.1B parameters per computation. It targets cost-efficient agent workloads, with stronger long-horizon planning, self-correction, and long-context processing.
- Trained across more than 10,000 real interaction environments for coding, task decomposition, and multi-source research.
- Alternates KDA linear attention and MLA layers at a 5:1 ratio to balance long-context efficiency and capability.
- Activates only 1/64 of its experts per token, reducing compute requirements versus its predecessor.
- Cluster-level caching reportedly cuts time-to-first-token by 60–80%+ on long inputs; the model is temporarily free on OpenRouter and planned for open-source release.
Sources (1)
蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-Flash
TL;DR - Ant Group’s Bailing team released Ling-3.0-Flash, a 124B-parameter hybrid-reasoning model that activates 5.1B parameters per computation. It targets cost-efficient agent workloads, with stronger long-horizon planning, self-correction, and long-context processing.
- Trained across more than 10,000 real interaction environments for coding, task decomposition, and multi-source research.
- Alternates KDA linear attention and MLA layers at a 5:1 ratio to balance long-context efficiency and capability.
- Activates only 1/64 of its experts per token, reducing compute requirements versus its predecessor.
- Cluster-level caching reportedly cuts time-to-first-token by 60–80%+ on long inputs; the model is temporarily free on OpenRouter and planned for open-source release.