Someone did this

Agents spend 10 million tokens and 85 minutes on one task. The best finishes 15% of them.

Reported by Zongxia Li and colleaguesJuly 9, 2026Not tested by benchr

What happened

Long-Horizon-Terminal-Bench, published on arXiv on July 9, 2026, grades agents on 46 terminal tasks across nine categories with partial credit rather than pass or fail. Fifteen frontier models were run. Against a 0.95 reward threshold the strongest model passes 15.2% of tasks first time, and 10.9% against a perfect score. The mean across all fifteen is 4.3% and 1.7%. A single task consumes on average 9.9 million tokens, about 231 episodes and 85.3 minutes of execution.

How it was done

  1. In46 terminal tasks across nine categories
  2. ThenGraded with dense partial credit, not pass or fail
  3. Then15 frontier models, ~231 episodes and 85 minutes per task
  4. OutBest 15.2% pass@1; mean across models 4.3%
Read the paper

What it cost

9.9 million tokens per task, on average, across fifteen models.

What this does not prove

  • A preprint, not peer reviewed.
  • 46 tasks is a small set; a few unusually hard ones move the mean a long way.
  • Scores are the authors' own runs, with the usual caveat that agent scaffolding differs between teams.

Someone did this · Not tested by benchr

Not reproducibleRecorded by benchr on

A preprint, and the authors' own runs. The interesting numbers here are the cost figures rather than the ranking, and those are less sensitive to how a benchmark is built.

Original source: arxiv.orgAlso covered by LHTB repository

Evidence

What each item establishes, and who established it.

  • Published by whoever did it

    What this shows The pass rates, the token and episode counts and the task breakdown are in the paper's own abstract and tables.

    Li et al., arXiv 2607.08964 · July 9, 2026arxiv.org

Where this leads