Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 27B MTP on v100 32GB: 54 t/s

Via r/LocalLlama
Wednesday, May 6, 2026 ยท 2:18AM
Summary

Just a quick note that I got a nice result using am17an's MTP branch of llama.cpp on v100 32GB SXM module using one of those pcie card adapters. Pulled and built in one shot, and llama-server ran without a hitch. Tested using am17an's MTP GGUF, q8_0 kv cache and 200k cache limit acting as vscode cop

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories