Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3x3090 test results

Via r/LocalLlama
Saturday, Aug 1, 2026 ยท 4:53PM
Summary

For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. CURRENT RESULTS: full moe offloading Prefill suffers 116 --> 72 t/s , generation 8-->14 t/s compared to previous case with no moe offlloading. ./llama-bench -m

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories