Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llama.cpp Gemma 4 using up all system RAM on larger prompts

Via r/LocalLlama
Monday, Apr 6, 2026 ยท 6:17AM
Summary

Something I'm noticing that I don't think I've noticed before. I've been testing out Gemma 4 31B with 32GB of VRAM and 64GB of DDR5. I can load up the UD_Q5_K_XL Unsloth quant with about 100k context with plenty of VRAM headroom, but what ends up killing me is sending a few prompts and the actual sy

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories