Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

MTP has no impact on my Qwen3.6 MoE performance

Via r/LocalLlama
Thursday, Jun 4, 2026 ยท 6:43AM
Summary

Hello I have an rtx 5060Ti and I tried running unsloth's Qwen3.6-35B GGUF with MTP. However in both cases I have around 60 tok/s. Here are my flags: llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --alias unsloth/Qwen3.6 --port 8002 --kv-unifie

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories