I've seen a few people in the comments on here and the other AI subs suggest mixing quantization for the KV cache to retain higher accuracy and still saving memory. I was running that for a while until I realized how wrong it is. I wrote a longer blogpost about it, but TL;DR is this benchmark run: m