🛰️ Daily AI Frontier
‹ back to 2026-08-27

像素级对标?智谱、Kimi 底层参数「撞衫」背后,藏着线性注意力的黄金窗口

Industry & News Efficiency & Systems

Ranking

Overall 68
Content 75
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 像素级对标?智谱、Kimi 底层参数「撞衫」背后,藏着线性注意力的黄金窗口

Merged summary

TL;DR - Zhipu’s GLM-5.3-Flash and Kimi K3 use highly similar KDA linear-attention configurations, including a gate_lower_bound of -5. The convergence highlights a narrow engineering sweet spot for stable, cost-efficient long-context inference rather than, by itself, evidence of copying.

  • KDA compresses history into a fixed-size state matrix, avoiding the growing KV-cache costs of conventional Softmax attention.
  • The -5 gate floor retains roughly 0.67% of old state, balancing numerical stability and long-term recall; higher values risk information buildup, while lower values risk excessive forgetting.
  • GLM-5.3-Flash combines KDA linear attention, DeepSeek-style sparse attention, and manifold-constrained hyper-connections to reduce inference costs.
  • Adding the gate constraint during later training may accelerate adoption of proven practices, but could cause subtle memory-distribution shifts in extreme long-context workloads.

Sources (1)

像素级对标?智谱、Kimi 底层参数「撞衫」背后,藏着线性注意力的黄金窗口

雷峰网 (AI科技评论) 2026-08-27
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:18:08.816961 UTC

TL;DR - Zhipu’s GLM-5.3-Flash and Kimi K3 use highly similar KDA linear-attention configurations, including a gate_lower_bound of -5. The convergence highlights a narrow engineering sweet spot for stable, cost-efficient long-context inference rather than, by itself, evidence of copying.

  • KDA compresses history into a fixed-size state matrix, avoiding the growing KV-cache costs of conventional Softmax attention.
  • The -5 gate floor retains roughly 0.67% of old state, balancing numerical stability and long-term recall; higher values risk information buildup, while lower values risk excessive forgetting.
  • GLM-5.3-Flash combines KDA linear attention, DeepSeek-style sparse attention, and manifold-constrained hyper-connections to reduce inference costs.
  • Adding the gate constraint during later training may accelerate adoption of proven practices, but could cause subtle memory-distribution shifts in extreme long-context workloads.
item →