Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I tested freshly merged DFlash in llama.cpp on Qwen 3.6 27B Local AI win. 4.44x faster at 36K context. Here are my findings RTX 6000 PRO.

Via r/LocalLlama
Tuesday, Jul 7, 2026 · 4:40PM
Summary

Hey guys, A month ago I posted my MTP benchmarks here (3.34x on Gemma 4). DFlash support just merged into llama.cpp (PR #22105), so I ran it on the same rig with the Qwen 3.6 27B and it beat my best MTP numbers at every draft length. DFlash is speculative decoding with a block diffusion drafter from

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories