Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Apple M5 isn't making full use of its matmul cores yet

Via r/LocalLlama
Thursday, Jul 23, 2026 · 4:28PM
Summary

At the moment MLX (and Llama.cpp for Macs) run 16bit activations everywhere. Despite this, the M5 generation silicon actually does support INT8 activations - it actually allows w4a8 d_type. It's just that no inference backends are using them yet I built some w8a8 kernels and have managed to get 1.4x

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories