Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 35B-A3B @ Q4 or Gemma 4 12B @ Q8?

Via r/LocalLlama
Sunday, Jun 14, 2026 ยท 9:30PM
Summary

Wondering how much model quantization matters here. Daily driver on my 32gb unified memory setup is the qwen model outputting ~15 tokens a second. Heard good things about the 12B Gemma 4 model so interested in trying it against my codebase. Given its size I can very comfortably fit the Q8 in. Hell,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
China may have accessed Mythos
The Verge AI · Industry & Money
Welcome to the AGI era of AI governance
Interconnects · Models & Research
Introducing the OpenAI Partner Network
OpenAI Blog · Models & Research
Back to all stories