Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Slow performance Unsloth Gemma 12B Q8

Via r/LocalLlama
Monday, Jun 29, 2026 · 9:54AM
Summary

I recently replaced GPT-OSS 20B Q4 with Gemma 4 12B Q8 but i went from roughly 70 t/s to 10 t/s. Am I doing something wrong? In the current session I am trying a Q5 modell with no change in performance meassured against the Q8. [Service] Type=simple User=root WorkingDirectory=/root/llama.cpp ExecSta

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories