Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)

Via r/LocalLlama
Tuesday, Jun 16, 2026 · 12:19AM
Summary

Link to last post Before anything else, I'd like to sincerely thank u/jipok_ for helping out by highlighting a few weak questions, categories and scoring issues, which have now been addressed (Dropping >100 questions, tuning the scoring methodology for more accuracy, etc). Thanks!! TL;DR: HalBench i

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories