How this site sources information

Where the data on benchr comes from and how it is kept current.

Pricing data

Per-token pricing for closed-source models comes directly from each provider's official pricing page: Anthropic, OpenAI, Google, Mistral, DeepSeek. Pricing for open-weight models hosted on third-party inference providers references the inference provider's own published rates when relevant. Where pricing changes, the article is updated and a changelog entry is added.

Benchmark scores

Benchmark numbers are sourced from the benchmark maintainers' published leaderboards. SWE-bench Verified scores come from swebench.com. LMSYS Arena scores come from lmarena.ai. ARC-AGI scores come from arcprize.org. When a provider publishes a model's score on a benchmark before it appears on the official leaderboard, the provider's published figure is used with attribution.

Capability ratings

Where this site assigns capability ratings (coding, reasoning, writing, vision, long context, multilingual) on a 0–100 scale, they are editorial estimates. They synthesize cited public benchmarks, provider documentation, and published third-party reporting. They are not measurements made by benchr, independently reproduced benchmark scores, or provider facts. A precise-looking rating remains an editorial planning aid, not a laboratory result.

Editorial estimates vs sourced figures (the tools)

The interactive tools — the recommender, calculator, charts, and benchmark explorer — all read one file, assets/data/models.json, and that file keeps a hard line between two kinds of number:

  • Official specification: pricing, context window, maximum output, release dates, and availability are copied from a linked provider document. “Official” means provider-published, not independently verified by benchr.
  • Provider-reported or third-party benchmark: a named public evaluation with its publisher identified. Provider-reported scores are not presented as independent results. If a comparable public score is unavailable, the benchmark field stays blank; benchr does not manufacture a replacement score.
  • Editorial judgment: the 0–100 capability profiles are a documented decision aid, not a benchmark. They synthesize cited public evidence and are labelled editorial wherever shown. The public dataset deliberately leaves latency, first-token, and tokens-per-second fields blank unless a reproducible measurement or a clearly attributed provider figure is available; a universal speed number would hide material variation by region, model version, request shape, and provider load.
  • Illustrative calculation: cost examples produced from stated assumptions and a cited rate card. They explain how to calculate a workload, not what a real customer paid.

Tool rankings mix published facts with clearly identified editorial judgment. Use them to make a shortlist, not as an objective leaderboard. A capability that a model does not support can count as zero; a missing benchmark or speed measurement is shown as unavailable and is not silently converted to zero. Before production, repeat the comparison with your own prompts, region, hardware, and quality rubric.

Editorial responsibility and release checks

The benchr Editorial Team is the accountable publication-level author. Its documented roles cover scope, evidence checks, Arabic editing, and release checks; the role labels do not claim separate employees. Before a new or materially revised article is released, material facts are checked against the cited evidence, unsupported fields stay unavailable, and the visible byline and Article schema are checked against the same public author profile.

AI tools may assist with organization, draft language, translation checks, code, and illustrations. They are not evidence and do not serve as the author or an independent reviewer. Source selection, factual approval, wording, and corrections remain editorial responsibilities.

What this site is and is not

This is an editorial publication that synthesizes public information. It is not a benchmarking lab. Current articles do not claim unpublished benchr test runs, private API-cost totals, or first-person time-on-tool results. Workload verdicts are editorial interpretations of cited benchmarks, official rate cards and specifications, and attributed public reporting. A verdict is not itself a measured result.

benchr does publish an original, CC BY 4.0 use-case evaluation package. It contains six bilingual suites, 30 fixed prompts, and scoring rubrics so readers can run a consistent evaluation. It deliberately contains no model outputs, scores, rankings, or claim that benchr executed the suites. A published test design is evaluation infrastructure, not a measured result.

For editorial calls, qualitative judgments (“better documented for long-document analysis,” “dialect performance needs local validation”) are preferred to unsupported precision. Each factual number should identify its provenance in the surrounding text or references. Estimated values must say that they are editorial estimates; hypothetical workload math must state its assumptions. If a comparison cannot be supported by a citable public source, it is described as an open question rather than as a private experiment.

If benchr publishes first-party testing in the future, that page will name the tester, date, model and software versions, hardware or region, prompts or dataset, scoring rubric, run count, and material limitations, with enough raw examples to audit the conclusion. Until those elements are available, the site does not present a scenario as a benchr test.

Update cadence

Pricing tables are checked against provider documentation when articles are revised. Model release and deprecation events are recorded after their official announcement and source review; benchr publishes no guaranteed turnaround. The schedule for systematic re-verification of all model data is “before major article revisions”. There is no fixed weekly or monthly cycle.

Corrections and disputes

If you find a number, date, or attribution that does not match the primary source, send a note to corrections@benchr.org. Material corrections are noted on the corrections page and in the article changelog.