Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.5-27B on RTX 5090 served via vLLM @ 77 tps

Via r/LocalLlama
Tuesday, Apr 21, 2026 ยท 12:44AM
Summary

After maxing out my cursor $20 sub and zai $10 sub for this month, I have resorted to a local llm setup. Got good outcome on RTX5090 running Qwen3.5 27B and achieved very good tps. Context window at 218k. It can even run 2 concurrent sessions with this config although per session speed drops as expe

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories