Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GGUF (llama.cpp) vs MLX Round 2: Your feedback tested, two models, five runtimes. Ollama adds overhead. My conclusion. Thoughts?

Via r/LocalLlama
Thursday, Mar 26, 2026 · 2:50PM
Summary

Two weeks ago I posted here that MLX was slower than GGUF on my M1 Max. You gave feedback, pointed out I picked possibly the worst model for MLX. Broken prompt caching (mlx-lm#903), hybrid attention MLX can't optimize, bf16 on a chip that doesn't do bf16. So I went and tested almost all of your hint

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories