TL;DR it's pretty goddamned fast; 69 tps decode at near max (256k) context with MTP on. prefill numbers went down to 893 at max context with prompt cache turned off. https://jdkruzr.github.io/3080bench/ here's how the tests were run: https://github.com/jdkruzr/3080bench/ there is probably more perfo