Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

EvoTensile: Evolutionary algorithms for AMD Tensile GEMM kernel tuning

Via r/LocalLlama
Friday, Jun 19, 2026 · 7:39AM
Summary

There has been an effort to tune kernels in hipBLASLt so the most basic matmuls can run faster. It's known that on Strix Halo (gfx1151), GEMM with NN and TN input layouts (used in inference) are already well-tuned, while NT and TT layouts (used in training) are not yet tuned. The tool we use to tune

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories