Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running Qwen3 30B A3B at 50 tok/s on RTX 5060 Ti

Via r/LocalLlama
Saturday, Jul 11, 2026 ยท 8:29AM
Summary

Experimented with some custom CUDA and C++ code that can now run a Qwen3-30B-A3B at 50-54 tok/s at float 8 on an RTX 5060 Ti with only 16 GB of VRAM. This speed is roughly 50% improvement to llama.cpp which runs at around 33-34 tok/s (with n-cpu-moe). These speedups come mostly from combining SOTA s

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories