Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

mistral.rs v0.9.0: up to 1.8x faster CPU decode than llama.cpp on x86 and ARM!

Via r/LocalLlama
Tuesday, Jul 7, 2026 · 3:36PM
Summary

https://preview.redd.it/nuk5rxceptbh1.png?width=1448&format=png&auto=webp&s=300344dd4c6552379e8536b81ba288be3d6dca3f On Qwen3 4B Q4_K, mistral.rs decodes faster than llama.cpp at every context depth we measured, on x86 (Sapphire Rapids) and ARM (GB10). We optimized mistral.rs at granular levels to a

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories