Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

9070xt inference for q3 qwen 27B

Via r/LocalLlama
Saturday, May 9, 2026 · 4:56PM
Summary

In llamacpp I'm getting 12tok/s, does this number look right to you and what can I do to increase this number (if possible)? cd ~/llama.cpp && ./build/bin/llama-server -m models/qwen-3.6-27b-abliterated-q3.gguf -ngl 999 -c 65536 (i need this, shrinking this is not an option) -np 1 -b 512 --ubatch-si

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories