Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GLM 5.2 and ik_llama.ccp

Via r/LocalLlama
Sunday, Jul 26, 2026 · 5:18AM
Summary

Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on the GPU. Works great at 8k context — rock solid, ~3.7 tok/s gen. Problem: the second I raise context

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories