Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

Via r/LocalLlama
Monday, Jun 8, 2026 · 3:24PM
Summary

Hey fellow Llamas, your time is precious, so I'll keep it short (while trying to explain everything lol). TL;DR: 33-35B MoE on a 16 GB GPU. Qwen3.6 35B-A3B: 13.3 GiB (was ~20.5). Laguna XS.2 33B-A3B: 14.6 GiB (was 18.8). Both measured on an RTX 3090, both under 16 GiB. Only the active experts stay o

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories