Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Linux - Why does llama.cpp ROCm consume SO much VRAM for KV cache compared to Vulkan?

Via r/LocalLlama
Thursday, May 14, 2026 · 6:13PM
Summary

I have a docker stack with a bunch of AI services and llama.cpp server is the brain. I've got a working vulkan yml snippet for llama.cpp but out of curiosity, I flipped it to ROCM (latest build) and did not see ANY performance improvement. In fact, I noticed that for the SAME model, SAME context set

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories