Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Industry & Money Story
Industry & Money

OpenAI finds roughly 30 percent of popular AI coding test is broken

Via The Decoder
Thursday, Jul 9, 2026 · 1:23PM
Summary

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark. The article OpenAI finds roughly 30 percent of popular AI coding test is broken appear

Continue reading the full article
Read at The Decoder
the-decoder.com
Back to all stories