In the recent kv rotation PR it was found that the existing q8 kv quants tank performance on AIME25, but can be recovered mostly with rotation
Via r/LocalLlama
Sunday, Mar 29, 2026 · 5:57PM
Summary
The comment: https://github.com/ggml-org/llama.cpp/pull/21038#issuecomment-4150413357 I think this could be great for existing q8 users. Personally I'll be sticking with fp16 for the foreseeable future.