Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen3.6-35B-A3B-APEX / 128K ctx on RTX 3060 12GB — 37 t/s gen with 72k ctx filled, PPL 3.25, offloading 17GB model

Via r/LocalLlama
Thursday, May 28, 2026 · 11:12AM
Summary

I'm posting this because it may be helpful to squeeze the 12GB VRAM in the 3060. All credit goes to spiritbuun's fork (github.com/spiritbuun/buun-llama-cpp) and mudler's APEX quantizations (huggingface.co/mudler). Spiritbuun's CUDA optimizations for NVIDIA GPUs — fused MMA fix, TurboQuant, fattn imp

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories