Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Should I be seeing more of a performance leap when using NVFP4, INT4, FP8 with VLLM over MXFP4, Q4, and Q8 with llama.cpp based inference on Blackwell based GPUs?

Via r/LocalLlama
Saturday, Apr 18, 2026 · 2:33PM
Summary

I hope I am doing something wrong here but I am seeing about almost double the t/s using LM studio with Qwen3.5 and Nemotron models than I am seeing with Nvidia’s own vLLM containers built for Spark. I was surprised I was only getting 15-ish t/s with Nemotron Nano NVFP4 in VLLM with Nvidia’s recomme

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories