Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

24+ tok/s from ~30B MoE models on an old GTX 1080 (8 GB VRAM, 128k context)

Via r/LocalLlama
Wednesday, May 13, 2026 · 8:41PM
Summary

I got Qwen 3.6 35B-A3B and Gemma 4 26B-A4B running on a $200 secondhand machine (i7-6700 / GTX 1080 / 32 GB RAM) using llama.cpp (the TurboQuant/RotorQuant KV cache quantisation allows 128k context within the 8 GB VRAM). Results (Q4_K_M models, 128k context): Model tok/s Key flags Qwen 3.6 35B-A3B ~

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories