Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 27b MTP vLLM

Via r/LocalLlama
Saturday, May 2, 2026 ยท 9:29AM
Summary

Hello everyone, i am banging my head trying to properly configure qwen 3.6 27b mtp in vllm. I am using vllm v0.20.0 in docker, unquantized model with tp4 (4 3090s), max context length. At low context size, mtp with value of 3 gives the best results: 48-50 tps generation speed. However, once the cont

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories