Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How can I get better response time by caching my system prompt?

Via r/LocalLlama
Monday, Jun 29, 2026 ยท 11:39AM
Summary

Hi, I've spent some time trying to find a solution to make my local AI cache the system prompt (unless it is already caching and hitting a wall on every new session is a thing)... I'm using Ornith 35b, with llama.cpp, on a Strix Halo (WIN10). It works great so far with my PI agent. I have around 7.1

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories