Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]

Via r/MachineLearning
Tuesday, Jun 2, 2026 · 8:38AM
Summary

I built CVE-Bench: 20 real-world CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, others), 5 frontier models, 3 prompt conditions, 300 runs total. Each agent runs in a sandboxed container and is scored against a hidden test_security.py derived from the maintainer's own fix. Binary

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories