Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-27B with MTP grafted on Unsloth UD XL: 2.5x throughput via unmerged llama.cpp PR

Via r/LocalLlama
Wednesday, May 6, 2026 · 11:45AM
Summary

Hey everyone, I've been working on getting Multi-Token Prediction (MTP) working with quantized GGUFs for Qwen3-27B and the results are pretty impressive. Here's what I put together: https://huggingface.co/havenoammo/Qwen3.6-27B-MTP-UD-GGUF These are Unsloth's UD XL quantizations of Qwen3-27B with th

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories