Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

RDNA2 flash attention isn’t enabled stock, I enabled it with this build and doubled my speed

Via r/LocalLlama
Tuesday, May 19, 2026 · 4:24AM
Summary

What's good everybody, I probably have the fastest possible setup on these AMD Radeon RDNA2 GPUs for one reason only. A custom binary that bypasses some assert statement causing a crash in today’s stock releases. This binary bypasses that assert and enables flash attention. Works for rocm lamma cpp

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories