Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM

Via r/LocalLlama
Tuesday, Aug 4, 2026 · 5:52PM
Summary

A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected experts in VRAM while the cold experts continue running on the CPU. The author’s results on Qwen3.6-35

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories