Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I wrote a fused MoE dispatch kernel in pure Triton that beats Megablocks on Mixtral and DeepSeek at inference batch sizes

Via r/LocalLlama
Sunday, Apr 5, 2026 · 6:03PM
Summary

Been working on custom Triton kernels for LLM inference for a while. My latest project: a fused MoE dispatch pipeline that handles the full forward pass in 5 kernel launches instead of 24+ in the naive approach. Results on Mixtral-8x7B (A100): Tokens vs PyTorch vs Megablocks 32 4.9x 131% 128 5.8x 12

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories