Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Idea for how to run GLM2 at a decent quant, need critique/feedback

Via r/LocalLlama
Monday, Jun 22, 2026 · 7:57PM
Summary

I am currently running a 4x 5060 ti P2P rig (64 GB VRAM total)where each card is running at gen 3 with 4 pcie lanes per card. My use case is inference only. During my benchmarking the bottleneck was compute, not pcie bandwidth for low concurrency inference tasks, such as a single user use case. This

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories