Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Is there a way to mitigate performance as context grows?

Via r/LocalLlama
Sunday, Apr 26, 2026 · 5:08PM
Summary

In my local LLM setup I get from 30 to 80 t/s generation at the beginning, but it drops quite a lot as context grows. I use llama.cpp/Vulkan with an MI50 and a V100, is there some command line flags that can improve this issue? Or some good practice other than restart the chat after some time?

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories