Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Llama.cpp's auto fit works much better than I expected

Via r/LocalLlama
Tuesday, Apr 21, 2026 · 6:07PM
Summary

I always thought with 32GB of VRAM, the biggest models I could run were around 20GB, like Qwen3.5 27B Q4 or Q6. I had an impression that everything had to fit in VRAM or I'd get 2 t/s. Man was I wrong. I just tested Qwen3.6 Q8 with 256k context on llama.cpp, with `--fit` on, the weights alone are bi

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories