Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-27B at 72 tok/s on RTX 3090 on Windows using native vLLM (no WSL, no Docker), portable launcher and installer

Via r/LocalLlama
Saturday, May 2, 2026 ยท 8:12AM
Summary

The angle here is native Windows, no WSL. Simple installation, open source, no telemetry. Not selling or promoting anything: https://github.com/devnen/qwen3.6-windows-server Numbers (RTX 3090, Windows 10): - 72 tok/s short prompt - 64.5 tok/s long prompt (~25k tokens) - 53.4 tok/s at 127k ctx (singl

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories