Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Expert Upcycling: Growing MoE capacity mid-training without increasing inference cost (7B→13B, ~32% GPU hours saved)

Via r/LocalLlama
Thursday, Apr 23, 2026 · 10:53PM
Summary

Author here, sharing a preprint we recently released. We're actively looking for feedback from this community before we revise. Motivation. Training large MoEs from scratch is expensive. All expert weights, gradients, and optimizer states must reside in accelerator memory regardless of how few are a

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories