Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings

Via r/LocalLlama
Saturday, Jul 4, 2026 · 3:11PM
Summary

Spent a while tuning llama.cpp for Qwen3.6 27B on a 9800X3D / 64GB / 5090 box and wanted to share the real distribution instead of just a headline number, since averages hide a lot. Ran with q8 KV cache, 192k context, MTP draft=10, spec-draft-p-min=0.5, batch/ubatch 512. Logged 6,454 samples across

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories