Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.

Via r/LocalLlama
Tuesday, Jun 9, 2026 · 1:50AM
Summary

MODS: if you wanna remove for slop, that's cool - once I have something like a llama.cpp patch I'll repost something people can use. I'll write a full article on my Medium account with how it works. Just got excited and wanted to share. *** Forgive the claude summary, in the readme, but the base wor

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories