Bartowski's DeepSeek-V4-Flash-MXFP4 GGUF, llama.cpp build 9851 (0eca4d490), deepseek4 arch. Ran the same n_ctx = 10240, same n_ubatch = n_batch = 8192, flash attention on — only difference is -ctk/-ctv: Cache type Total KV cache (CUDA0) CUDA0 compute buffer f16 (default, no -ctk/-ctv set) ~425 MiB 1