Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - SinkProbe tests attention sinks and position-dependent recall at million-token context lengths. Across four controlled small models, it finds that sinks arise from the training objective rather than architecture, while gating fails to reproduce previously reported improvements at this scale.
- SinkProbe measures sink mass, massive activations, position-resolved recall, and the recency gap.
- The four evaluated models differ only in how they mix information across tokens and depth.
- Attention gating did not reproduce its previously published reduction in first-token attention.
- Sink mass, activation magnitude, and positional bias varied independently, so no single metric captures long-context behavior.
Sources (1)
Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - SinkProbe tests attention sinks and position-dependent recall at million-token context lengths. Across four controlled small models, it finds that sinks arise from the training objective rather than architecture, while gating fails to reproduce previously reported improvements at this scale.
- SinkProbe measures sink mass, massive activations, position-resolved recall, and the recency gap.
- The four evaluated models differ only in how they mix information across tokens and depth.
- Attention gating did not reproduce its previously published reduction in first-token attention.
- Sink mass, activation magnitude, and positional bias varied independently, so no single metric captures long-context behavior.