Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ggml-webgpu: Improve prefill speeds for k-quants + refactor matmul for Q4/Q5/Q8 and k-quants by yomaytk · Pull Request #24225 · ggml-org/llama.cpp

Via r/LocalLlama
Tuesday, Jun 9, 2026 · 2:41AM
Summary

This PR improves matmul performance for k-quants. The following table shows the improvement on the pp512 test in M2 pro. quant model master (t/s) PR (t/s) speedup Q2_K qwen3 0.6B Q2_K - Medium 817.86 ± 6.14 1991.81 ± 6.87 2.44x Q3_K qwen35 4B Q3_K - Medium 92.54 ± 0.13 302.24 ± 0.37 3.27x gemma4 E4B

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories