Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llama.cpp DeepSeek v4 Flash experimental inference

Via r/LocalLlama
Sunday, Apr 26, 2026 · 10:20AM
Summary

Hi, here you can find experimental llama.cpp support for DeepSeek v4, and here there is the GGUF you can use to run the inference with "just" (lol) 128GB of RAM. The model, even quantized at 2 bit, looks very solid in my limited testing, and the speed of 17 t/s in my MacBook M3 Max is quite interest

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories