Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I did some model hacks, and got GLM5.2 from about 2.5 tok/s to >50 tok/s on my GH200 system.

Via r/LocalLlama
Wednesday, Jun 24, 2026 · 1:30PM
Summary

G'day. This is part 3 on my Local LLM adventures. I have a crazy system hacked server-to-desktop system: Component Spec GPUs 2x Hopper H100, 96 GB HBM3 each CPUs 2x Grace, 72 cores each Host memory 480 GB LPDDR5X per Grace, 960 GB total So I can run technically run GLM5.2. Except the naive settings

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories