🛰️ Daily AI Frontier
‹ back to 2026-09-09

Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

arXiv cs.CL LLMs & Foundation Models Sara Rizwan, Samaanah Abdus Salam 2026-09-08

TL;DR - SinkProbe tests attention sinks and position-dependent recall at million-token context lengths. Across four controlled small models, it finds that sinks arise from the training objective rather than architecture, while gating fails to reproduce previously reported improvements at this scale.

  • SinkProbe measures sink mass, massive activations, position-resolved recall, and the recency gap.
  • The four evaluated models differ only in how they mix information across tokens and depth.
  • Attention gating did not reproduce its previously published reduction in first-token attention.
  • Sink mass, activation magnitude, and positional bias varied independently, so no single metric captures long-context behavior.

view merged work →