Someone did this

On the hardest web tasks, people finish 10%. The best agent finishes 8%.

Reported by Zejun Xu and colleaguesAugust 9, 2026Not tested by benchr

What happened

CAP, published on arXiv on August 9, 2026, is a benchmark of 420 cross-site tasks on 108 real websites across 24 domains, built around interactions that need real visual understanding rather than form-filling. Full success rates: Manus 8.0%, Fellou 7.0%, Comet 6.0%, Genspark 5.0%, Claude-4.5-Sonnet 5.0%, Dia 4.0%, DeepSeek-V4-Flash 2.9%, GPT-5 2.0% — against a human baseline of 10.0%. On partial completion the gap opens up: humans 35.0%, Comet 48.0%, Manus 23.0%.

How it was done

  1. In420 tasks, 108 real websites, 24 domains
  2. ThenSplit into complex actions and complex perception
  3. ThenEight agents and a human baseline run against the same set
  4. OutPerception, not manipulation, is the dominant failure
Read the paper

What this does not prove

  • A preprint, not a peer-reviewed paper.
  • Scores were measured against live websites and will drift as those sites change.
  • The human baseline is a small comparison group under the same time pressure, not a measure of what a determined person could do given all day.

Someone did this · Not tested by benchr

Not reproducibleRecorded by benchr on

A preprint. The agent scores are the authors' own runs on live websites, so they carry the usual caveat that a site can change under a benchmark.

Original source: arxiv.org

Where this leads