Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

TurboQuant on MLX: 4.6x KV cache compression with custom Metal kernels (Qwen 32B at 98% FP16 speed)

Via r/LocalLlama
Saturday, Mar 28, 2026 · 9:07AM
Summary

Implemented TurboQuant (Google's new KV cache compression paper) for MLX with fused Metal kernels. Results on Qwen2.5-32B, M4 Pro 48GB: - 4.6x compression, 0.98x FP16 speed, identical quality - 16K context: 4.2GB cache → 897MB The main challenge was speed — went from 0.28x to 0.98x FP16 through fuse

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories