Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

*Lower* generation speed with H100 and H200 than with RTX 5090?

Via r/LocalLlama
Monday, Jun 15, 2026 · 9:44AM
Summary

I've tried renting some cloud instances to get an idea of the speed of various GPUs. I'm using a recent version llama.cpp with CUDA 12.8 support. I've tried running a 31B dense model, Q6, on an RTX 5090 and an H100, and the results surprised me. The 5090 generates at about 57 tok/sec, while the H100

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories