Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-35B-A3B Q4 262k context on 8GB 3070 Ti = +30tps

Via r/LocalLlama
Friday, May 22, 2026 · 10:11PM
Summary

..and on 8GB VRAM I can even push the context to 320K, 400K, 512K, and yes.. 1M. But it does start to slow down noticeably beyond 150k so I'd only do this if I ever really want the larger context. This is using APEX-I-Quality or Q4_K_XL quants both are better than Q4_K_M (IQ4_NL_XL for beyond 512k c

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories