Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Single 3090 with Q4 Qwen 27B, context dropped from 137k to 14k with MTP enabled. Is it normal?

Via r/LocalLlama
Wednesday, May 27, 2026 ยท 12:27AM
Summary

Note: Latest version of llama.cpp (b4c0549a49be9e6dc59ac9d0a5bc21dbda910774) My run command: ``` llama-server \ --temp 0.6 \ --top-p 0.95 \ --top-k 20 \ --presence_penalty 0.0 \ --min-p 0.00 \ --gpu-layers all \ -m /home/eleung/huggingface/unsloth/Qwen3.6-27B-MTP-GGUF/Qwen3.6-27B-UD-Q4_K_XL.gguf \ -

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories