Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ggml-zendnn : add Q8_0 quantization support by z-sachin · Pull Request #23414 · ggml-org/llama.cpp

Via r/LocalLlama
Wednesday, Jul 15, 2026 · 12:23PM
Summary

Benchmark Results Benchmark configuration: threads = 96 type_k = bf16 type_v = bf16 Llama-3.1-8B-Instruct Q8_0 Prompt Size GGML_CPU_Q8_0 t/s ZenDNN_Q8_0 t/s Gain 256 472.28 730.87 54.75% 512 450.86 832.48 84.64% 768 446.81 864.52 93.49% 1024 439.58 800.15 82.03% 2048 405.07 778.34 92.15% tg128 33.08

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories