Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

BeeLlama v0.2.0 – major DFlash update. Single RTX 3090: Qwen 3.6 27B up to 164 tps (4.40x), Gemma 4 31B up to 177.8 tps (4.93x). Prompt processing speed near baseline.

Via r/LocalLlama
Friday, May 22, 2026 · 5:34PM
Summary

BeeLlama v0.2.0 is here! Not quite a pegasus, but close enough. GitHub | Qwen 3.6 27B Quick Start | Gemma 4 31B Quick Start Full Gemma 4 31B support with efficient DFlash implementation and vision. Major Qwen 3.6 27B performance update from lower DFlash overhead, cleaner prefill handling, drafter K/

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories