🛰️ Daily AI Frontier
‹ back to 2026-07-24

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

arXiv cs.LG Efficiency & Systems Alagappan Valliappan 2026-07-23

TL;DR - Windowed-MTP limits speculative-decoding draft attention to a sliding window while retaining full-context target verification. At million-token context, it reduces decode-step cost by 28–44% without changing the target model’s verified output distribution.

  • Bounds draft KV usage to a constant-size working set, dropping about 99% of draft KV entries at 1M context.
  • Uses a training-free, drop-in sliding window with an attention sink only for the MTP draft.
  • Prevents full-context draft attention from dominating—or negating—speculative-decoding gains at long contexts.
  • A compact ring buffer reclaims draft KV representing 7.7–11% of total KV at 1M context without acceptance or quality loss.

view merged work →