Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

GPoUr with ~12gb vram and a 3080 getting 40tg/s on qwen3.6 35BA3B w/ 260k ctx

Via r/LocalLlama
Thursday, Apr 16, 2026 ยท 8:23PM
Summary

The TheTom's turboquant's GPU accelerated turboquant (turbo3) has unlocked high context gains for the 35BA3B family. I can now achieve ~40tg/s via the following GPU-POOR compilation flags and configuration: cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_CUDA_F16=ON -DGGML_CUDA_FOR

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories