Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Why are quants on KV cache increase before weight quants?

Via r/LocalLlama
Tuesday, Jun 2, 2026 ยท 10:11PM
Summary

I'm cases where ram is limited I've seen a preference for increasing kvcache precision instead of the weight precision. I.e. 8bit kvcache but only 4bit weights. But I can't seem to find a solid explanation as to why?

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories