TL;DR: Full GLM-5.2 (753B MoE) quantized to Int4-Int8Mix + NVFP4 4-bit KV cache, TP=4 across 4× DGX Spark (GB10) at 100K context, run on Terminal-Bench 2.1 with the same agent scaffold (Terminus-2) as the official numbers. Result: 63/89 = 70.8% vs the official full-precision 81.0%. Caveat up front: