Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

RTX 5070 Ti 16GB + 32GB RAM: Running Qwen3.6-35B-A3B Q8_0 @ 44 t/s (128K context)

Via r/LocalLlama
Friday, Apr 24, 2026 · 11:02AM
Summary

32GB DDR5 RAM. unsloth/Qwen3.6-35B-A3B-GGUF Q8_0 : 36.9 GB LM studio settings: - GPU Offload: 40 - Offload MoE Experts to CPU: 26 -Try mmap: on -K cache:Q8_0 -V cache:Q8_0 llama.cpp will be better.

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories