🛰️ Daily AI Frontier
‹ back to 2026-08-13

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

arXiv cs.CL LLMs & Foundation Models Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo 2026-08-12
Representative image for Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

TL;DR - This study identifies architecture-specific massive activation patterns in hybrid linear-attention LLMs. The findings link these patterns to full-attention placement and activation-cancellation timing, clarifying how hybrid architectures approach standard attention behavior.

  • Massive activations spike immediately before full-attention layers and may persist across intervening linear-attention layers.
  • Denser full attention increasingly connects these spikes into the stable activation pattern seen in full-attention LLMs.
  • The patterns recur across five architectures, six hybrid configurations, five data domains, and models from 1.2B to 397B parameters.
  • Full-attention output gating substantially reduces activation magnitude without removing its layerwise organization.

view merged work →