Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Fused MoE dispatch kernel in pure Triton: 89-131% of Megablocks, runs on AMD with zero code changes

Via r/LocalLlama
Wednesday, May 27, 2026 ยท 12:58PM
Summary

I've been working on MoE inference and wrote a fused dispatch kernel entirely in Triton, no CUDA. At inference batch sizes (up to 512 tokens) it reaches 89-131% of Megablocks(Stanford's CUDA-optimized MoE lib), and the same kernel runs on AMD MI300X with no changes. Mixtral-8x7B on A100. The biggest

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories