Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Skipping 90% of KV dequant work → +22.8% decode at 32K (llama.cpp, TurboQuant)

Via r/LocalLlama
Friday, Mar 27, 2026 · 2:56PM
Summary

I’ve been working on an open source TurboQuant implementation for KV cache compression in llama.cpp and ran into a hard bottleneck: dequantization. At long context (32K on M5 Max), dequant alone was taking around 40 percent of decode time. I tried fixing it the usual way: - register LUTs - SIMD tric

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories