Ran identical benchmarks on both 16” MacBook Pros with 40 GPU cores and 128GB unified memory across three Qwen 3.5 models (122B-A10B MoE, 35B-A3B MoE, 27B dense) using oMLX v0.2.23. Quick numbers at pp1024/tg128: 35B-A3B: 134.5 vs 80.3 tg tok/s (1.7x) 122B-A10B: 65.3 vs 46.1 tg tok/s (1.4x) 27B dens