Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6 27B on dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k context working

Via r/LocalLlama
Wednesday, Apr 29, 2026 · 8:40AM
Summary

I’ve been testing Qwen3.6 27B on a pretty non-standard local setup and figured the numbers might be useful for anyone looking at the newer 16GB Blackwell cards. Hardware: 2x RTX 5060 Ti 16GB 32GB total VRAM Proxmox LXC 16 vCPU ~60GB RAM CUDA 13 / Torch 2.11 nightly vLLM nightly: 0.19.2rc1.dev Model:

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories