Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GPU VRAM only for small models with llama.cpp: is it possible?

Via r/LocalLlama
Sunday, May 24, 2026 · 3:02PM
Summary

I'm still in my learning process and so far I've been able to make satisfying use of my setup (4070 with 12GB VRAM + 32GB RAM and iGPU for my GUI). I've been able to run both Gemma4 26B and Qwen 3.6 35B MoEs up to high quants with large context and have about 40 t/s with both. However, I'd like to t

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories