Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Can't get over 250TPS on RTX5090 with Qwen3.5-4B

Via r/LocalLlama
Saturday, May 30, 2026 · 1:33PM
Summary

My main model is qwen3.6-27b-mtp and I'm getting around 100tps and 2500tps prefill, which is great. I've tried adding a second small model for auxiliary tasks, and even when it's the only model running, it doesn't go over 200-250tps. I'm building llama.cpp and running on docker windows. I've also tr

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories