Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I benchmarked 42 STT models on medical audio with a new Medical WER metric — the leaderboard completely reshuffled

Via r/LocalLlama
Thursday, Apr 9, 2026 · 4:00PM
Summary

TL;DR: I updated my medical speech-to-text benchmark to 42 models (up from 31 in v3) and added a new metric: Medical WER (M-WER). Standard WER treats every word equally. In medical audio, that makes little sense — “yeah” and “amoxicillin” do not carry the same importance. So for v4 I re-scored the b

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories