Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Can't replicate Reddit numbers with Qwen 27B on a 3090TI.

Via r/LocalLlama
Thursday, Apr 30, 2026 · 11:26AM
Summary

I feel like i'm going insane. I see people here posting 30 - 100+ tok/s (100+ being with speculative decoding) on a 3090 with Qwen 3.6 27B. I'm trying to replicate this but my performance numbers are nowhere near that. I have tried llama.cpp with Unsloth's Q4XL and Q4_K_M GGUF's. On that i got like

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories