Choose, test, price, and manage AI models
Every tool reads the same sourced model record. Start with the decision in front of you, then carry the shortlist into the next step.
Capability map
What can AI actually do?
Start from the job rather than the model. Every record names its source, its limits, and the date benchr last read the vendor documentation behind it.
What AI can actually do
New capabilities, recently rechecked records, quick wins, and the ones that recently stopped working.
Open Discover → LedgerThe capability record
Steps, prompts, limits, sources and a source-check history for each concrete thing a provider documents.
Read the ledger → Goal firstTasks and workflows
Say what you want done and see the documented routes, ranked by effort, with the multi-step chains behind the bigger jobs.
Browse by task →Decision suite
What do you need to decide?
Explore builds the shortlist, Labs tests it against your cases, and Migration Assistant keeps production integrations from aging silently.
Model workspace
Check the evidence, filter the index, save a shortlist, and share the current view.
Explore models → EvaluateModel and prompt test
Bring cases from Prompt Workbench, hide model names, add cost and latency, and export the results.
Run a test → OperateMigration assistant
Look up a retired API ID, check the documented replacement, compare prices, and build a checklist.
Plan a migration →Calculator
Cost Calculator
Enter your token volumes and workload mix. Get a monthly cost estimate per model with caching and batch discounts applied.
Tokens
Token Counter
Paste a prompt to count tokens as you type — exact for OpenAI through the real o200k_base tokenizer, labeled estimates for Claude and Gemini — then compare the cost across every priced model.
Prompt QA
Prompt Workbench
Build an explicit prompt contract, capture version A, inspect whether one or several sections changed, and send private A/B cases to Labs.
Recommender
Model Recommender
Answer a few questions about your task, budget, and quality bar. Get a ranked shortlist of models that fit your requirements.
Compare
Side-by-side Compare
Pick any two or more of the 38 indexed models. Pricing, benchmarks, context windows, and capability ratings side by side.
Charts
Benchmark Charts
Capability or official benchmark plotted against price, with the value frontier drawn in: the models nothing cheaper beats, and what every other model is beaten by.
Tracker
Model Tracker
Current status for announced, available, and deprecated models, with the dates and migration notes needed to plan an integration.
Timeline
Release Timeline
Chronological history of tracked model releases across OpenAI, Anthropic, Google, Meta, Mistral, and open-weight providers.
Developer reference
AI model identifier finder
Resolve a model name, slug, API ID, or repository ID, then copy the recorded value with its source and verification date.
Data access
Developer data
Public read endpoints, model and lifecycle CSV files, an RSS feed, and a downloadable retirement calendar.
Local inference
Local AI reference
Inspect source-linked open-weight records and plan a conservative Q4 memory fit before downloading a model or buying hardware.
Local controls
Settings & local data
Choose the site theme and typeface, inspect benchr browser storage, create a bounded backup, or clear selected local categories.
Also useful
- Pricing index → — Input, output, caching, and batch rates for all 38 models.
- Rankings → — Models ranked by the benchr Rating composite score.
- Local AI planner → — Source-linked open-weight records, Q4 memory planning, and an honest “what can I run?” starting point.
- Price per use case → — Cheapest model for chat, coding, RAG, agents, and classification.
- How to reduce token usage → — Caching, batching, and routing tactics.
- Provider hubs → — OpenAI, Anthropic, Google Gemini, and open-weight records.
When each tool is useful
Start with Rankings for a broad shortlist. Use Compare when two or three candidates differ across price, context, and benchmarks. Move to the Calculator once you know your token mix; cache rate and output length can change the bill more than the headline input price.
Charts help you test different benchmark priorities. The tracker records releases and retirements. Together, they make it easier to revisit a decision when prices, scores, or model status change.
Data policy
Tool data comes from the shared benchr model index and is reviewed against provider documentation. Pricing, context windows, release dates, and model IDs are treated as factual fields. Composite ratings and some benchmark fills are editorial estimates where providers do not publish directly comparable numbers; those estimates are documented in Methodology.