Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[P] Fused MoE Dispatch in Pure Triton: Beating CUDA-Optimized Megablocks at Inference Batch Sizes

Via r/MachineLearning
Sunday, Apr 5, 2026 · 6:07PM
Summary

I built a fused MoE dispatch kernel in pure Triton that handles the full forward pass for Mixture-of-Experts models. No CUDA, no vendor-specific code. On Mixtral-8x7B (A100), it beats Stanford's Megablocks at inference-relevant batch sizes (131% at 32 tokens, 124% at 128 tokens). At larger batches M

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories