We now running 2x models with vLLM (MiniMax M2.7 quantized into MXFP4_16 (iq4_nl)) on 8xR9700 and we use Qwen3-27b-BF16 on 4x7900xtx. also at vLLM. We use it locally not for large demand, but fully offline inference. What we can fit in 400GB VRAM via llama cpp, and maybe there is someone here who ru