Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

unsloth vs bartowski MTP ggufs

Via r/LocalLlama
Monday, Jun 1, 2026 ยท 8:32AM
Summary

I noticed that bartowski's MTP ggufs are bigger than unsloth. I asked bartowski and he said he used Q8_0 quant for the MTP head. So I compare the decoding performance of the two. /build/bin/llama-server -m ~/gguf/Qwen3.5-4B-Q4_0.gguf --host 0.0.0.0 --port 8080 -c 4096 -fa on --no-mmap -np 1 -ngl 99

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories