Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 q8 at 50t/s or q4 at 112 t/s?

Via r/LocalLlama
Friday, Apr 17, 2026 · 5:26PM
Summary

What are some ways that you would go about thinking about choosing between the two for use in a harness like pi? Did a good bit with q4 yesterday and it was so consistent and reliable I had it set to 131k context and it worked through 2 compactings on a clearly defined task without messing the whole

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories