Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

VLLM gives 5x speed of llama but quants not available (unsloth/gguf). What to do?

Via r/LocalLlama
Thursday, May 28, 2026 · 2:58PM
Summary

Hi - I want to run unsloth dynamic quant on vllm. Why? vllm is giving faster prefill speed - Llama - i get 800-1000 tokens/sec - Vllm - i get 5k-10K tokens/sec Tried using Qwen3.6-35B-A3B FP8 official. Machine is RTX A6000 - ampere 48gb Unsloth q8 quant (on llama testing) gives correct pandas code,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories