Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model

Via r/LocalLlama
Wednesday, Jul 8, 2026 · 8:08PM
Summary

TL;DR: Full GLM-5.2 (753B MoE) quantized to Int4-Int8Mix + NVFP4 4-bit KV cache, TP=4 across 4× DGX Spark (GB10) at 100K context, run on Terminal-Bench 2.1 with the same agent scaffold (Terminus-2) as the official numbers. Result: 63/89 = 70.8% vs the official full-precision 81.0%. Caveat up front:

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories