Someone did this
On the hardest web tasks, people finish 10%. The best agent finishes 8%.
Reported by Zejun Xu and colleaguesAugust 9, 2026Not tested by benchr
What happened
CAP, published on arXiv on August 9, 2026, is a benchmark of 420 cross-site tasks on 108 real websites across 24 domains, built around interactions that need real visual understanding rather than form-filling. Full success rates: Manus 8.0%, Fellou 7.0%, Comet 6.0%, Genspark 5.0%, Claude-4.5-Sonnet 5.0%, Dia 4.0%, DeepSeek-V4-Flash 2.9%, GPT-5 2.0% — against a human baseline of 10.0%. On partial completion the gap opens up: humans 35.0%, Comet 48.0%, Manus 23.0%.
How it was done
- In420 tasks, 108 real websites, 24 domains
- ThenSplit into complex actions and complex perception
- ThenEight agents and a human baseline run against the same set
- OutPerception, not manipulation, is the dominant failure
What this does not prove
- A preprint, not a peer-reviewed paper.
- Scores were measured against live websites and will drift as those sites change.
- The human baseline is a small comparison group under the same time pressure, not a measure of what a determined person could do given all day.
Someone did this · Not tested by benchr
Not reproducibleRecorded by benchr on
A preprint. The agent scores are the authors' own runs on live websites, so they carry the usual caveat that a site can change under a benchmark.