Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Worse quality with MTP - Qwen 3.6, Gemma 4

Via r/LocalLlama
Thursday, Jun 25, 2026 · 7:10AM
Summary

Hi. I am self-hosting Qwen 3.6 27B Q8_K_XL with Llama.cpp on 4x5070ti. (All 4 cards are on single x16 slot bifurcated to 4x4 with risers). I've been testing it on several work repos with Opencode CLI and in like 8/10 situations the output of non-MTP model is far superior to the MTP ones. The prompt

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories