Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ggml-cuda: add flash-attn support for DKQ=320/DV=256 with ncols2=32 (… by lnigam · Pull Request #22286 · ggml-org/llama.cpp

Via r/LocalLlama
Tuesday, Apr 28, 2026 · 9:48PM
Summary

Improves the speed of Mistral Small 4 on CUDA (there was a CPU fallback before) (I wonder if it’s somehow related to the upcoming Mistral model? Maybe not)

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories