Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

Via r/LocalLlama
Saturday, May 9, 2026 · 4:05PM
Summary

TL;DR New llama.cpp fork! I wanted a Windows-friendly inference to run Qwen 3.6 27B Q5 on a single RTX 3090 with speculative decoding, high context without excess quantization, and vision enabled. No option did this out of the box for me without VRAM and/or tooling issues (this was before MTP PR for

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories