Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek-V4-Flash W4A16+FP8 with MTP self-speculation: 85 tok/s @ 524k on 2× RTX PRO 6000 Max-Q

Via r/LocalLlama
Sunday, May 10, 2026 · 6:22PM
Summary

TL;DR: DeepSeek-V4-Flash running at 85.52 tok/s @ 524k ctx and ~111 tok/s @ 128k single-stream on 2× RTX PRO 6000 Max-Q pasta-paul's DeepSeek-V4-Flash-W4A16-FP8 quant is great, but its MTP head silently gets stripped at load time (HF transformers has it in _keys_to_ignore_on_load_unexpected), so --s

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories