Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 35B GGUF: NTP vs MTP quantization results across GPUs and CPUs

Via r/LocalLlama
Wednesday, May 20, 2026 · 3:42PM
Summary

Hey r/LocalLLaMA, We’ve released our ByteShape Qwen 3.6 35B GGUF quantizations in two families: standard NTP (Next Token Prediction or non-MTP) and MTP. Blog / Download NTP Models / Download MTP Models TL;DR For NTP, “pick the largest quant that fits” worked surprisingly well. Lower bpw was not auto

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories