Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ICYM: llama.cpp b9455 --SM Tensor KV Cache Fix is MERGED

Via r/LocalLlama
Monday, Jun 1, 2026 · 8:08PM
Summary

Them boys can cook, one big fix after another! If you're running --sm tensor on multi-gpu this is the KV cache quantization fix https://github.com/ggml-org/llama.cpp/releases/tag/b9455 JohannesGaesslercommented5 days ago This PR implements support for the combination of -sm tensor and quantized KV c

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories