Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

BeeLlama v0.3.1 – latest llama.cpp with extras! DFlash, MTP, q6_0 cache, TurboQuant. Single RTX 3090: Qwen 3.6 27B & Gemma 4 31B up to 177.8 tps (4.93x over baseline)

Via r/LocalLlama
Thursday, Jun 4, 2026 · 9:25PM
Summary

BeeLlama v0.3.0 and v0.3.1 are here! Big architectural update to align the fork with upstream llama.cpp and integrate all its additions like MTP and Gemma 4 12B support, while also updating DFlash to handle complex configurations like multi-slot and multi-GPU. Now also recommended by club-3090! Than

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories