Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-MTP-27B on Tesla V100 @ 55 TPS (llama.cpp) — Any way to push this higher without quality loss?

Via r/LocalLlama
Wednesday, Jun 10, 2026 · 10:36AM
Summary

Hey everyone, I'm running Qwen3.6-MTP-27B-MTP (Q4_K_M) with llama.cpp server on a Tesla V100, and I'm currently getting around 55 tokens/sec. I'm trying to find out whether there are any configuration changes that could increase throughput further without reducing output quality. 55 TPS seems lower

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories