Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Testing llama.cpp MTP support on Qwen3.6 - RTX 5090

Via r/LocalLlama
Sunday, May 17, 2026 · 6:00AM
Summary

Setup: - RTX 5090, 32 GB, Linux - Built llama.cpp from 4f13cb7 (the official ghcr.io/ggml-org/llama.cpp:server-cuda image hasn't picked up the merge yet as of writing — had to docker build from source with CUDA_DOCKER_ARCH=120) - Unsloth's Qwen3.6-27B-MTP-GGUF Q5_K_M and Qwen3.6-35B-A3B-MTP-GGUF UD-

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories