Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s)

Via r/LocalLlama
Thursday, Jul 16, 2026 · 6:47PM
Summary

so ive been messing around with qwen3.6 35b a3b (MXFP4 gguf) on my 3060 12gb, doing the usual cpu/gpu offload thing where half the expert layers sit in ram and get pulled over pcie whenever needed. and like everyone whos done this knows the gpu just sits there idle waiting for experts to show up, pc

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories