Interactive tools
Ten connected tools for discovering, testing, comparing, pricing, and migrating AI models. The suite reads one source-backed data layer, so a shortlist can move from research to evaluation without starting over.
Decision suite
Start with a decision, not a directory
Explore builds the shortlist, Labs tests it against your cases, and Migration Assistant keeps production integrations from aging silently.
Model intelligence platform
Search evidence, filter the catalog, compare economics, save collections and watchlists, then share the state.
Open Explore → EvaluateModel evaluation lab
Bring prompts and outputs, score them blind with a rubric, include cost and latency, and export the report.
Run an eval → OperateMigration assistant
Look up retired API IDs, inspect replacements, compare price and benchmark deltas, and generate a checklist.
Plan a migration →Calculator
Cost Calculator
Enter your token volumes and workload mix. Get a monthly cost estimate per model with caching and batch discounts applied.
Tokens
Token Counter
Paste any prompt for a live token count — exact for OpenAI via the real o200k_base tokenizer, labeled estimates for Claude and Gemini — plus the cost across every priced model.
Recommender
Model Recommender
Answer a few questions about your task, budget, and quality bar. Get a ranked shortlist of models that fit your requirements.
Compare
Side-by-side Compare
Pick any two or more of the 29 indexed models. Pricing, benchmarks, context windows, and capability ratings side by side.
Charts
Benchmark Charts
Visual comparison of SWE-bench, MMLU, GPQA, and pricing across the full model index. Sortable and filterable.
Tracker
Model Tracker
Live status of announced, available, and deprecated models. Useful for keeping your integration plans current.
Timeline
Release Timeline
Chronological history of every major model release across OpenAI, Anthropic, Google, Meta, Mistral, and the open-weight field.
Also useful
- Pricing index → — Input, output, caching, and batch rates for all 29 models.
- Rankings → — Models ranked by the benchr Rating composite score.
- Price per use case → — Cheapest model for chat, coding, RAG, agents, and classification.
- How to reduce token usage → — Caching, batching, and routing tactics.
When each tool is useful
Start with Rankings when you need a broad shortlist. Move to Compare when the decision is between two or three models and the trade-off is not obvious from a single benchmark. Use Calculator last, after you know your expected token mix, because small differences in cache rate and output length can change the bill more than the headline input price.
The chart and tracker pages are for maintenance work: checking whether a model still fits your benchmark target, watching release cadence, and spotting when a cheaper tier has caught up enough to replace a frontier model. The goal is not to crown one permanent winner, but to make model selection repeatable as prices and scores move.
Data policy
Tool data comes from the shared benchr model index and is reviewed against provider documentation. Pricing, context windows, release dates, and model IDs are treated as factual fields. Composite ratings and some benchmark fills are editorial estimates where providers do not publish directly comparable numbers; those estimates are documented in Methodology.