Them boys can cook, one big fix after another! If you're running --sm tensor on multi-gpu this is the KV cache quantization fix https://github.com/ggml-org/llama.cpp/releases/tag/b9455 JohannesGaesslercommented5 days ago This PR implements support for the combination of -sm tensor and quantized KV c