Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Via r/MachineLearning
Saturday, Aug 1, 2026 · 9:27AM
Summary

While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics. Also, clinically meani

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories