Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

RTX 5080 16GB: Qwen3.6 35B MoE at 128k context — 56 tok/s, and why MTP doesn't help

Via r/LocalLlama
Wednesday, May 20, 2026 · 11:33AM
Summary

MTP (Multi-Token Prediction) just merged into mainline llama.cpp at b9190. I promised u/WarthogConfident4039 a Qwen3.6 benchmarking round. Three configs, tested at real coding-agent context lengths (not just 512 tokens). The main finding surprised me. TL;DR: 35B Q4_K_XL, no MTP, --fit-target 1536**,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories