Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Built LazyMoE — run 120B LLMs on 8GB RAM with no GPU using lazy expert loading + TurboQuant

Via r/LocalLlama
Sunday, Apr 12, 2026 · 7:53PM
Summary

I'm a master's student in Germany and I got obsessed with one question: can you run a model that's "too big" for your hardware? After weeks of experimenting I combined three techniques — lazy MoE expert loading, TurboQuant KV compression, and SSD streaming — into a working system. Here's what it loo

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories