Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running Qwen3.6-35B-A3B on a laptop RTX 4060 (8GB) — what worked, what didn't, and a surprising speculative-decoding result

Via r/LocalLlama
Friday, Jun 5, 2026 · 8:25PM
Summary

TL;DR: I spent a long session tuning a 35B MoE on a tiny 8GB laptop GPU. Three things mattered a lot (--no-mmap, VRAM headroom, closing CPU-hungry apps). Several "obvious" optimizations did nothing because of this model's hybrid architecture (TurboQuant, Flash Attention, even i-quants made it worse)

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories