Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How to run Qwen3.5-27B with speculative decoding with llama.cpp llama-server?

Via r/LocalLlama
Monday, Apr 13, 2026 · 11:39AM
Summary

I run it on 2xRTX 3090. This is part of my llama-server presets file: [Qwen3.5-27B-bartowski] load-on-startup = true alias = Qwen3.5-27B-bartowski hf = bartowski/Qwen_Qwen3.5-27B-GGUF:Q8_0 hfd = bartowski/Qwen_Qwen3.5-2B-GGUF:Q8_0 draft-min = 1 draft-max = 4 temp = 0.6 top-p = 0.95 top-k = 20 min-p

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories