Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DeepSeek-V4-Flash 284B on 5.3GB of memory

Via r/LocalLlama
Sunday, Aug 2, 2026 · 7:28AM
Summary

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE models activate a few B params per token, so keep the shared core and KV cache resident and stream the s

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories