Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Got MTP + TurboQuant running — Qwen3.6-27B -- 80+ t/s at 262K context on a single RTX 4090

Via r/LocalLlama
Friday, May 8, 2026 · 9:15PM
Summary

So I've been messing around trying to get MTP working alongside TBQ4_0 (TurboQuant's lossless 4.25 bpv KV cache) on Qwen3.6-27B for my own use. So after a day of vibecoding I think I may have gotten something viable. Went from about 43 t/s when I first got it compiling to 80-87 t/s after optimizing.

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories