Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I stopped trusting model benchmarks and started running my own eval set, here is what changed[D]

Via r/MachineLearning
Thursday, Jun 25, 2026 · 9:22AM
Summary

Three things broke my faith in published benchmarks recently. One, Kimi K2.7 Code shipped with plus 21.8 percent on Kimi Code Bench v2, plus 11 percent on Program Bench, plus 31.5 percent on MLS Bench Lite. All three are Moonshot's own benchmarks. None were submitted to DeepSWE, which is the one ind

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories