Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Attention Drift: What Autoregressive Speculative Decoding Models Learn

Via r/LocalLlama
Tuesday, May 12, 2026 · 7:10PM
Summary

Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturbation and long-context inputs. We identify a previously-unreported phenomenon we call \textbf{attention drift}: as the drafter generates successive to

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories