Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ROCm vs Vulkan vs vLLM on Dual R9700's

Via r/LocalLlama
Sunday, Jun 21, 2026 · 1:53PM
Summary

Just wanted to share these numbers I saw running Qwen3.6 35BA3 and Qwen3.6 27B and the big increase I saw going to vLLM. I was just expecting better concurrency but ended up with a lot better speeds. llama.cpp services Running ROCm and Vulkan Model Backend Gen 35B-A3B Q6_K_XL (MTP) ROCm ~106 t/s 27B

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories