Im running the raw version straight from the minimax release on hugging face (https://huggingface.co/MiniMaxAI/MiniMax-M2.7) on 3 rtx pro 6000's on vllm. So no quantization. And i'm not going to lie something feels off about it. Same workloads in our coding environment, including our re-usable evals