Someone did this
Agents spend 10 million tokens and 85 minutes on one task. The best finishes 15% of them.
Reported by Zongxia Li and colleaguesJuly 9, 2026Not tested by benchr
What happened
Long-Horizon-Terminal-Bench, published on arXiv on July 9, 2026, grades agents on 46 terminal tasks across nine categories with partial credit rather than pass or fail. Fifteen frontier models were run. Against a 0.95 reward threshold the strongest model passes 15.2% of tasks first time, and 10.9% against a perfect score. The mean across all fifteen is 4.3% and 1.7%. A single task consumes on average 9.9 million tokens, about 231 episodes and 85.3 minutes of execution.
How it was done
- In46 terminal tasks across nine categories
- ThenGraded with dense partial credit, not pass or fail
- Then15 frontier models, ~231 episodes and 85 minutes per task
- OutBest 15.2% pass@1; mean across models 4.3%
What it cost
9.9 million tokens per task, on average, across fifteen models.
What this does not prove
- A preprint, not peer reviewed.
- 46 tasks is a small set; a few unusually hard ones move the mean a long way.
- Scores are the authors' own runs, with the usual caveat that agent scaffolding differs between teams.
Someone did this · Not tested by benchr
Not reproducibleRecorded by benchr on
A preprint, and the authors' own runs. The interesting numbers here are the cost figures rather than the ranking, and those are less sensitive to how a benchmark is built.
Original source: arxiv.orgAlso covered by LHTB repository
Evidence
What each item establishes, and who established it.
- Published by whoever did it
What this shows The pass rates, the token and episode counts and the task breakdown are in the paper's own abstract and tables.
Li et al., arXiv 2607.08964 · July 9, 2026arxiv.org