Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

FP4 inference in llama.cpp (NVFP4) and ik_llama.cpp (MXFP4) landed - Finally

Via r/LocalLlama
Saturday, Apr 25, 2026 · 3:42PM
Summary

Both llama.cpp and ik_llama.cpp now have FP4 support — but with different flavors worth knowing about. llama.cpp recently merged NVFP4 (Nvidia's block-scaled FP4, `GGML_TYPE_NVFP4 = 40`), with CUDA kernels landing in `mmq.cuh`, `mmvq.cu`, `convert.cu` and others. ik_llama.cpp has had MXFP4 (`GGML_TY

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories