Best AI News β€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.5-122B-Q5-MTP - Qwen3.5-122B-Q6-MTP

Via r/LocalLlama
Saturday, May 16, 2026 Β· 9:54PM
Summary

for anyone who cares... πŸ˜„ prompt = spen a 1000 tokens unsloth MTP models strix halo llama.cpp:server-rocm-mtp \ --spec-type draft-mtp \ --spec-draft-n-max 3 Qwen3.5-122B-Q5-MTP-General n_decoded = 100 tg = 29.77 t/s n_decoded = 179 tg = 27.95 t/s n_decoded = 254 tg = 26.80 t/s n_decoded = 4056 tg =

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories