Have there been any notable difference between Q8 and FP16 on both the weights and the cache? I know the jump to Q8 is significant. I would test myself, but FP16 on my setup is painfully slow. Also side question, is ~14TPS around the number I should be expecting on a Strix Halo running 3.6 27B at Q8