Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Benchmark scores vs arena.ai performance [D]

Via r/MachineLearning
Tuesday, May 12, 2026 ยท 6:54PM
Summary

If you take two open-source models: Gemma4-31B and Qwen3.5-27B, you might notice that Qwen beats Gemma on almost all benchmark sets. On the other hand, the situation is reversed on arena.ai: Gemma beats Qwen quite decisively, by about 50 ELO points in different categories. Is there a good explanatio

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories