Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 35B-A3B is very usable with 12GB of VRAM

Via r/LocalLlama
Friday, May 8, 2026 · 9:22PM
Summary

Hardware: RTX 3060 12GB 32GB DDR4-3200 Windows CUDA 13.x Model: Qwen3.6-35B-A3B-MTP-IQ4_XS.gguf The model is a 35B MoE, so -ncmoe matters a lot. Lower -ncmoe means more MoE blocks stay on GPU. Main takeaway 12GB VRAM feels like a very practical size for this model. It lets you keep enough MoE blocks

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories