Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6:35b UD Q4_K_M 80 tok/s on Nvidia P40

Via r/LocalLlama
Wednesday, Jul 15, 2026 · 11:10PM
Summary

https://preview.redd.it/nrxkobib5hdh1.png?width=1147&format=png&auto=webp&s=a36a282635bcbb23c0c2bdffa855689eab5e76f9 I'm running Unsloth's Qwen3.6:35B UD Q4_K_M with a 100k context on a single NVIDIA P40 (24GB) using TheTom's TurboQuant fork of llama.cpp. I know this is only burst speed, but I thoug

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories