Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GLM 5.2, what speeds are we getting locally?

Via r/LocalLlama
Saturday, Jun 20, 2026 · 8:11PM
Summary

Can everyone that is able to run GLM 5.2 locally report what their inference engine, system specs, quantization, context size, and tokens/sec? If you're getting great numbers expect follow-up questions. I'll start: llamma.cpp, 6x RTX 3090, 128 DDR5, i7-13700K, unsloth UD-IQ2_M, 90K context @ Q8_0 KV

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories