Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request [P]

Via r/MachineLearning
Friday, May 29, 2026 · 8:54AM
Summary

We built a monokernel that runs the full decode sequence as one GPU-resident program on AMD MI300X, with some neat optimizations. The die topology is central to the result, we map memory access patterns to the physical layout, compute units group by their associated IOD, and the hardware runs at its

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories