Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp

Via r/LocalLlama
Wednesday, Jul 8, 2026 · 5:49PM
Summary

TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to ~400. Dropping DCP raises prefil to ~750. I'm running 4 drafted tokens vs Z.ai's rec of 5. Decode is heavily dependent on prose. Thinking gets ~20 tok/s. Co

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories