Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6-35B-A3B KV cache bench: f16 vs q8_0 vs turbo3 vs turbo4 from 0 to 1M context on M5 Max

Via r/LocalLlama
Tuesday, Apr 28, 2026 · 5:17PM
Summary

Took TheTom's TurboQuant Metal fork of llama.cpp (github.com/TheTom/llama-cpp-turboquant, the feature/turboquant-kv-cache branch) and ran a depth sweep on Qwen 3.6-35B-A3B Q8. TheTom had already published M5 Max numbers up to 32K. I wanted to see what the curves looked like once you push them. Hardw

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories