Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6:27B VRAM 16GB 5080: MTP Quant, Speeds, and Configs

Via r/LocalLlama
Tuesday, May 19, 2026 · 8:14PM
Summary

For those of you running Qwen3.6:27B on 16GB VRAM, what quantization did you settle on? For my primary purpose as a HA voice assistant, I've found my ideal target to be >50 tg and >800 pp. Qwen3.5:9B works really fast, but I'm experimenting with higher intelligence. Offloaded the vision model to CPU

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories