Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig

Via r/LocalLlama
Sunday, Aug 2, 2026 · 1:05AM
Summary

Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/TacoTakumi/DeepSeek-V4-Flash-0731-GGUF I requantized only the 129 routed expert tensors of DeepSeek-V4

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories