Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek-V4-Flash-0731 on Bosgame M5 with RTX PRO 6000 Max-Q eGPU

Via r/LocalLlama
Saturday, Aug 1, 2026 · 9:27PM
Summary

Here are my numbers: Quant Size Layout Decode Prefill Draft acceptance UD-Q8_K_XL 150.8 GiB 20 layers CUDA0 / 23 ROCm0 + drafter 44.0 t/s 564 t/s 0.535 UD-Q4_K_XL 144.4 GiB 22 / 21 + drafter 48.4 t/s 585 t/s 0.532 UD-Q2_K_XL 90.2 GiB entirely on CUDA0, no drafter 59.5 t/s 1513 t/s — I let claude por

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories