Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Question: Llama cpp, whats good right now for: MTP, KV cache quant, Long context.

Via r/LocalLlama
Thursday, May 28, 2026 · 7:21AM
Summary

Used the vllm version of https://github.com/noonghunna/club-3090 It worked fine for myabe 20 40k context, havent tried the new one. Anyone used the new llama.cpp patched one for single 3090? The project is starting to seem very bloated, at least readme wise. I use https://github.com/Indras-Mirror/ll

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories