Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Hot Experts in your VRAM! Dynamic expert cache in llama.cpp for 27% faster CPU +GPU token generation with Qwen3.5-122B-A10B compared to layer-based single-GPU partial offload

Via r/LocalLlama
Wednesday, Apr 15, 2026 · 3:20AM
Summary

Claude cooked on the code, but I wrote this post myself, caveman style. I wanted to play with Qwen3.5-122B, but I don't have a unified memory system to work with, and 15 tok/s was rough. 23 tok/s is still rough but honestly noticeably faster when streaming responses. Tl;dr: We keep track of which ex

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories