ResearchPod Summary
As LLMs scale to million-token contexts, speculative decoding—which uses a draft head to propose tokens for a target model to verify—faces a "draft-attention tax." The paper investigates why built-in Multi-Token-Prediction (MTP) heads, which typically use full-attention over the entire KV cache, become a performance bottleneck at long context. The author seeks a way to accelerate these drafts without sacrificing the output quality or requiring expensive retraining.
Instead of forcing the draft head to attend to the entire million-token history, the author applies a StreamingLLM-style sliding window combined with an attention sink. This restricts the draft's attention to a fixed, small number of recent tokens plus initial sink tokens. Because the target model (which performs the actual verification) still maintains full-attention over the entire context, the process remains mathematically lossless regarding the final output distribution. The technique is drop-in, requires no training, and allows for the physical reclamation of the unused draft KV cache into a compact ring buffer.
This work addresses a critical scaling failure in modern frontier models. As models move toward hybrid or linear-attention targets to save memory, the draft head's full-attention read becomes the dominant cost of the decode step. Windowed-MTP effectively "decouples" the draft's performance from the context length, ensuring that speculative decoding remains a viable acceleration strategy even at the million-token regime.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.