Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How to use llama.cpp to quantize to NVFP4?

Via r/LocalLlama
Wednesday, Jun 3, 2026 · 2:27AM
Summary

Trying to run MiniMax M2.7 NVFP4 via llama.cpp but not seeing any GGUFs anywhere on huggingface. So I’m guessing I would need to quantize to NVFP4.GGUF myself. Is this possible with llama.cpp, and if so, what commands need to be run to make this happen?

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories