Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Via r/LocalLlama
Saturday, May 9, 2026 ยท 11:57AM
Summary

Just wanted to share my config in hopes of helping other 12GB GPU owners achieve what I see as very respectable token generation speeds with modest VRAM. Using the latest llama.cpp build + MTP PR, I got over 80 tok/sec with 80%+ draft acceptance rate on the benchmark found here: https://gist.githubu

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories