Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpp

Via r/LocalLlama
Wednesday, Jul 1, 2026 · 12:01AM
Summary

Bartowski's DeepSeek-V4-Flash-MXFP4 GGUF, llama.cpp build 9851 (0eca4d490), deepseek4 arch. Ran the same n_ctx = 10240, same n_ubatch = n_batch = 8192, flash attention on — only difference is -ctk/-ctv: Cache type Total KV cache (CUDA0) CUDA0 compute buffer f16 (default, no -ctk/-ctv set) ~425 MiB 1

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories