Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

My biggest Issue with the Gemma-4 Models is the Massive KV Cache!!

Via r/LocalLlama
Friday, Apr 3, 2026 · 1:48PM
Summary

I mean, I have 40GB of Vram and I still cannot fit the entire Unsloth Gemma-4-31B-it-UD-Q8 (35GB) even at 2K context size unless I quantize KV to Q4 with 2K context size? WTF? For comparison, I can fit the entire UD-Q8 Qwen3.5-27B at full context without KV quantization! If I have to run a Q4 Gemma-

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories