Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I benchmarked 31 STT models on medical audio — VibeVoice 9B is the new open-source leader at 8.34% WER, but it's big and slow

Via r/LocalLlama
Friday, Mar 27, 2026 · 9:13AM
Summary

TL;DR: v3 of my medical speech-to-text benchmark. 31 models now (up from 26 in v2). Microsoft VibeVoice-ASR 9B takes the open-source crown at 8.34% WER, nearly matching Gemini 2.5 Pro (8.15%). But it's 9B params, needs ~18GB VRAM (ran it on an H100 since I had easy access, but an L4 or similar would

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories