Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Could someone please help explain these results?

Via r/LocalLlama
Monday, May 25, 2026 ยท 3:01AM
Summary

I'm running Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf on 12 GB VRAM and 32 GB RAM via the TurboQuant variant of llama.cpp. I increased the --n-cpu-moe value from 8 to 30, and my inference rate doubled! (17 to 34 tok/s). Shouldn't it have slowed down from the CPU having to do so much more work? Here is the com

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories