Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.5 27B at 1.1M tok/s on B200s, all configs on GitHub

Via r/LocalLlama
Thursday, Mar 26, 2026 · 7:49PM
Summary

Pushed Qwen 3.5 27B (the dense one, not MoE) to 1,103,941 tok/s on 12 nodes with 96 B200 GPUs using vLLM. 9,500 to 95K per node came from four changes: DP=8 over TP=8, context window from 131K to 4K, FP8 KV cache, and MTP-1 speculative decoding. That last one was the biggest -- without MTP, GPU util

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories