Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

ClawBench: Can AI Agents Complete Everyday Online Tasks? 153 tasks, 144 live websites, best model at 33.3% [R]

Via r/MachineLearning
Tuesday, Apr 14, 2026 · 5:21PM
Summary

We introduce ClawBench, a benchmark that evaluates AI browser agents on 153 real-world everyday tasks across 144 live websites. Unlike synthetic benchmarks, ClawBench tests agents on actual production platforms. Key findings: The best model (Claude Sonnet 4.6) achieves only 33.3% success rate GLM-5

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories