Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6 27B NVFP4 + MTP on a single RTX 5090: 200k context working in vLLM

Via r/LocalLlama
Wednesday, May 6, 2026 · 2:05PM
Summary

So I spent some time testing Qwen3.6 27B NVFP4 on my RTX 5090 and wanted to share the numbers, since most of the recent good posts are either around 48GB cards, FP8, or llama.cpp/GGUF. This is not a "best possible setup" claim. More like: this is what I got working, here are the exact params, here a

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories