Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

Via r/LocalLlama
Sunday, Aug 2, 2026 · 8:42AM
Summary

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quant of Qwen3.5-122B-A10B I'm aware of at any size — the 82 GiB build edges out 94–95 GiB 6-bit builds,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories