Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s!

Via r/LocalLlama
Wednesday, May 20, 2026 · 12:41AM
Summary

Hey r/DeepSeek, Who says we need an H100 cluster or the latest expensive GPUs to run frontier MoE models? I wanted to see how far we could push a single node of consumer legacy hardware, so we spent less than $2,500 total to build a budget machine that successfully runs DeepSeek-V4-Flash (284B total

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories