Just a quick observation for all three users of Jetson AGX Orin 64GB in this sub: q8_0 quant gives >20% faster prefill (prompt processing) than q6_k, and 10% faster than q4_k_xl. Tested with Unsloth Qwen3.6-27B-MTP-GGUF on recent llama.cpp build. I don't have statistics at hand, but from observation