Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Industry & Money Story
Industry & Money

New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously

Via The Decoder
Saturday, May 16, 2026 · 1:08PM
Summary

Researchers at Carnegie Mellon University built a new benchmark that measures how far AI agents can go when exploiting real vulnerabilities in Google's V8 engine. Mythos leads GPT-5.5 by a wide margin but costs twelve times as much. The article New benchmark shows Claude Mythos and GPT-5.5 can devel

Continue reading the full article
Read at The Decoder
the-decoder.com
Back to all stories