Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

Via r/MachineLearning
Wednesday, Jul 22, 2026 · 7:04AM
Summary

Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in Mixture-of-Experts (MoE) training. If you've trained MoEs, you know that optimizer state is usually t

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories