Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Looking for a working Deepseek-v4-Flash quant

Via r/LocalLlama
Wednesday, May 27, 2026 · 5:58PM
Summary

Best I tried so far is https://huggingface.co/nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF with the custom llama.cpp fork, but it suffers from low quality and random incoherent output. VLLM wouldn't support anything other than H100s for DS4. Any quantization out there that works on llama.cpp/vllm?

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories