Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GLM 5.1 Locally: 40tps, 2000+ pp/s

Via r/LocalLlama
Saturday, Apr 25, 2026 ยท 4:31PM
Summary

After some sglang patching and countless experiments, managed to get reap-ed nvfp4 version running stable and FAST on 4 x RTX 6000 Pros (limited to 350W). Very happy with performance and quality. Inference software is still under-optimized for those cards. I think we will see their true potential un

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories