I’m using llama-swap with llama.cpp. I mainly use opencode + pi.dev and I’m seeing frequent massive prompt reprocessing / prefills even tho the prompts are very similar between requests. Example behavior: context grows to +50k tokens LCP similarity often shows 0.99+ but sometimes n_past suddenly fal