Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

Via r/LocalLlama
Monday, Aug 3, 2026 · 8:25PM
Summary

Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bother with a 2018 server The model is 156 GB. That number decides everything before speed matters: Platf

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories