Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)

Via r/LocalLlama
Wednesday, Aug 5, 2026 · 8:15PM
Summary

daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20% faster e2e in reality, if at all, but hey, it's free and open! Apache 2.0

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories