Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Need help tuning cache in llama-server

Via r/LocalLlama
Sunday, Jul 12, 2026 · 7:16AM
Summary

Hey I am running a few models on a strix halo box. Especially for the larger models (like Qwen 3.5 122B) they work okayish performance wise if the cache is utilised properly but a full cache miss at 100k context causes roughly 10-20 minute of PP time - which is extremely annoying. I will first show

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories