Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

We added W8A8 activation quantization to MLX — prefill went from 2.84s to 2.52s on M5 Pro

Via r/LocalLlama
Monday, May 25, 2026 · 8:16AM
Summary

Hey, I work on inference tooling at Mininglamp AI. We needed faster prefill for a 4B VLM running on Apple Silicon. Problem was MLX only does weight-only quant — activations stay FP16 the whole way through. So we wrote Cider, a small SDK that adds W8A8 activation quant on top of MLX. Numbers on M5 Pr

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories