Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I benched quad 20GB 3080s on Vast AI for code generation with Qwen3.6-27B so you don't have to (it's even better than quad 5060Tis)

Via r/LocalLlama
Thursday, Jul 23, 2026 · 2:38AM
Summary

TL;DR it's pretty goddamned fast; 69 tps decode at near max (256k) context with MTP on. prefill numbers went down to 893 at max context with prompt cache turned off. https://jdkruzr.github.io/3080bench/ here's how the tests were run: https://github.com/jdkruzr/3080bench/ there is probably more perfo

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories