Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode

Via r/LocalLlama
Tuesday, Jul 14, 2026 · 8:24PM
Summary

GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode (generic - coding/structured) and memory remaining to run also a Mimo 2.5 in parallel for image/audio input, both tp 8. https://x.com/i/status/2077123292352204943

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories