Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Tesla P40 running qwen 3.6

Via r/LocalLlama
Monday, May 18, 2026 · 4:48PM
Summary

Does anyone know why qwen 3.6 MTP spec decoding won't work with Tesla P40 when the K cache is quantized? I was able to get mtp qwen 3.6 27B Q5 running at 20t/s on my tesla p40. But only after removing any quantization of the K cache (running at F16). I had no trouble running turbo3 k cache without M

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories