For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. CURRENT RESULTS: full moe offloading Prefill suffers 116 --> 72 t/s , generation 8-->14 t/s compared to previous case with no moe offlloading. ./llama-bench -m