Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I tracked a major cache reuse issue down to Qwen 3.5’s chat template

Via r/LocalLlama
Wednesday, Apr 8, 2026 · 5:51PM
Summary

Over the last week, I’ve been investigating cache misses while optimizing local agent workflows on my M5 Max. My setup used oMLX.ai as a backend with agents like OpenCode.ai and Pi.dev, but I reproduced the same behavior with other backends like llama.cpp too. At first, I assumed this was an inferen

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories