像素级对标?智谱、Kimi 底层参数「撞衫」背后,藏着线性注意力的黄金窗口
TL;DR - Zhipu’s GLM-5.3-Flash and Kimi K3 use highly similar KDA linear-attention configurations, including a gate_lower_bound of -5. The convergence highlights a narrow engineering sweet spot for stable, cost-efficient long-context inference rather than, by itself, evidence of copying.
- KDA compresses history into a fixed-size state matrix, avoiding the growing KV-cache costs of conventional Softmax attention.
- The -5 gate floor retains roughly 0.67% of old state, balancing numerical stability and long-term recall; higher values risk information buildup, while lower values risk excessive forgetting.
- GLM-5.3-Flash combines KDA linear attention, DeepSeek-style sparse attention, and manifold-constrained hyper-connections to reduce inference costs.
- Adding the gate constraint during later training may accelerate adoption of proven practices, but could cause subtle memory-distribution shifts in extreme long-context workloads.