Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-27B-INT4 clocking 100 tps with 256k context length on 1x RTX 5090 via vllm 0.19

Via r/LocalLlama
Sunday, Apr 26, 2026 ยท 8:37AM
Summary

Thanks to the community the Qwen3.6-27B speed keeps getting better. The following improves upon my recipe from yesterday and delivered a whopping 100+ tps (TG). Model: https://huggingface.co/Lorbus/Qwen3.6-27B-int4-AutoRound - MTP supported - KLD is decent (much better than NVFP4 per the linked post

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories