Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Gemma 4 31B QAT GGUF loads with MTP branch, but outputs repeated - any working recipe?

Via r/LocalLlama
Sunday, Jun 7, 2026 · 9:02AM
Summary

I’m trying to run: unsloth/gemma-4-31B-it-qat-GGUF gemma-4-31B-it-qat-UD-Q4_K_XL.gguf on an RTX 5090 32GB using llama.cpp Gemma 4 MTP PR branch. Main model loads. Without the MTP assistant head, /v1/chat/completions returns repeated <unused49>. I also tried the public MTP assistant head: boxwrench/g

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories