Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[D] 60% MatMul Performance Bug in cuBLAS on RTX 5090 [D]

Via r/MachineLearning
Friday, Apr 10, 2026 · 5:51PM
Summary

cuBLAS dispatches an inefficient kernel for every batched FP32 workload, from 256×256 to 8192×8192×8. It only uses ~40% of the available compute on RTX GPUs. Tested with RTX 5090, but likely all RTX non-Pro GPUs are affected. I tested with the latest CUDA 13.2.51, cuBLAS 13.3.0, and driver 595.58.03

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories