Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG

Via r/LocalLlama
Sunday, Aug 2, 2026 · 1:49AM
Summary

Hey y'all. I'll be concise. TL;DR: DS V4-Flash-0731 @ UD-IQ2_M running fully in VRAM on 3xMI50s (90.9 GB model, 96 GB VRAM). Actual speed on llama-server is: - Text Generation: ~15-16 tokens/second stable. Never dipped below 14 tokens/second, even when the model was spitting out a 30K token long rep

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories