Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GLM-5.2 (744B, 2-bit) at 7.3 tok/s on 4×3090 + 192GB — and why IQ1_M wasn't any faster

Via r/LocalLlama
Friday, Jun 19, 2026 · 12:06AM
Summary

TLDR: For the first time, I feel relief that they could shut down the cloud services and I would be ok. I got my 4th 3090 and then unsloth dropped the Q2 and Q1. I wrote nothing else here its from CC, so it might be wrong. GLM-5.2 UD-IQ2_M runs across 4×3090 + RAM expert offload at ~7.3 tok/s. Two d

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories