Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Is Attention sink without Positional Encoding unavoidable? [D]

Via r/MachineLearning
Thursday, Apr 30, 2026 · 9:20AM
Summary

TL;DR: As soon as I remove Positional Encoding (PE) from Self or Cross-attention, I start seeing vertical hot lines in attention heatmaps. Is there any way to make a model have query-conditioned attention without PE? So, I've been trying to pre-train a couple types of Transformer based models (small

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories