Why are quants on KV cache increase before weight quants?
Via r/LocalLlama
Tuesday, Jun 2, 2026 ยท 10:11PM
Summary
I'm cases where ram is limited I've seen a preference for increasing kvcache precision instead of the weight precision. I.e. 8bit kvcache but only 4bit weights. But I can't seem to find a solid explanation as to why?