Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Follow-up: GLM-5.2 NVFP4 on four DGX Sparks — the MTP mystery is solved, and it's now ~24 tok/s at 128K context

Via r/LocalLlama
Friday, Jul 3, 2026 · 6:33AM
Summary

Follow-up: GLM-5.2 NVFP4 on four DGX Sparks — the MTP mystery is solved, and it's now ~24 tok/s at 128K context This is a follow-up to my earlier post about running GLM-5.2 NVFP4 on 4x DGX Spark at 128K context. Short version of that post: 128K worked at ~15 tok/s with MTP1, and there was a painful

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories