Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

110 tok/s with 12GB VRAM on Qwen3.6 35B A3B and ik_llama.cpp

Via r/LocalLlama
Thursday, May 21, 2026 · 11:09AM
Summary

Had been getting great MTP performance with llama.cpp on my RTX 4070 Super 12GB, until they actually merged the MTP PR. Then, performance tanked and was barely above non-MTP. So, I decided to try out ik_llama.cpp since it also supports MTP and is apparently better optimized for CPU offloading. I did

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories