Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[D] MXFP8 GEMM: Up to 99% of cuBLAS performance using CUDA + PTX

Via r/MachineLearning
Monday, Mar 30, 2026 ยท 7:48AM
Summary

New blog post by Daniel Vega-Myhre (Meta/PyTorch) illustrating GEMM design for FP8, including deep-dives into all the constraints and design challenges introduced by MXFP8. Link: https://danielvegamyhre.github.io/2026/03/29/mxfp8-gemm.html Original Tweet: https://x.com/vega_myhre/status/203829361420

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories