Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA [R]

Via r/MachineLearning
Wednesday, May 27, 2026 · 9:25PM
Summary

New preprint. A Mixture-of-Experts inference kernel (TritonMoE) written entirely in OpenAI Triton, targeting portability across NVIDIA and AMD without vendor-specific code. Highlights: A fused gate+up GEMM computes both SwiGLU projections from shared tile loads, eliminating 35% of global memory traf

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories