Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Dynamic KV Cache Quantization and Load-on-demand mmproj/MTP: my llama.cpp wishlist

Via r/LocalLlama
Thursday, Jun 4, 2026 · 6:52PM
Summary

We all know the struggle of optimizing your VRAM usage: quantized model, quantized kvcache, mmproj off. I'm often frustrated by the tradeoffs I have to make in these areas. On my RTX 5090, I can fit: Qwen3.5-27B @ Q6_K Mmproj enabled, MTP off q8_0 kvcache 150k context That brings me to 29/32 GB. I c

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories