Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

2× Radeon AI PRO R9700 (RDNA4/gfx1201) on vLLM 0.22.1 — how we fixed the long-context decode cliff (and what we learned chasing FP8)

Via r/LocalLlama
Friday, Jun 19, 2026 · 6:14PM
Summary

Posting our setup for the (apparently growing) club of people running multiple R9700s on vLLM. Big shout-out to u/AustinM731 — their AITER Unified Attention post was the single most useful thing we found, and I want to (a) confirm it works, (b) share where our findings lined up vs differed, and (c)

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories