Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How do i prevent llama.cpp from offloading on Swap?

Via r/LocalLlama
Thursday, Jun 11, 2026 · 11:22AM
Summary

I have tried preventing this issue by using llama.cpp flags. However, I still have the issue: whenever I'm close to my 96GB of RAM, llama-server / llama.cpp decides to offload the KV cache onto my swap. This usually happens when I'm at 91-92GB of RAM and I still have 4GB to spare. Is there a more ag

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories