Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I can fit 28% more context after building llama.cpp with OpenBLAS. Huh?

Via r/LocalLlama
Thursday, Jun 4, 2026 · 4:58PM
Summary

I've noticed a weird difference when building llama.cpp with the Vulkan and OpenBLAS backends vs. building with the Vulkan backend only. It seems like llama.cpp can fit significantly more context in VRAM when built with OpenBLAS than when built without. I don't know if this is expected behavior, a b

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories