Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC

Via r/LocalLlama
Wednesday, Jul 22, 2026 · 5:30PM
Summary

GLM-5.2 UD-Q4_K_XL GGUF @ 12.2 tok/s output // 30.9 tok/s input on a real 10.7k-token document using llama.cpp RPC - At 10.7k context: 10.2 tok/s output with coherent long-form generation Two parallel requests: 14.5 tok/s aggregate Context: 2x 16,384-token slots Model size: 436 GiB Hardware: 16x AMD

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories