AI model charts
Compare capability with API price, see which models nothing cheaper beats, then set your own priorities. The frontier line and the cost-per-point column are computed from the pricing index, not chosen by hand.
Capability vs API price
Each dot is a model. The horizontal axis is the mean of input and output price per million tokens. For the vertical axis you can keep the 0–100 capability weighting (coding 40%, reasoning 40%, writing 20%) or switch to an official benchmark, in which case only the models whose provider published that number are plotted. The violet line is the value frontier: the models nothing cheaper beats. Pointer users can select a dot; keyboard and screen-reader users can use the table below.
Value frontier — no cheaper model scores higher
View the chart as a table, with cost per point
Loading chart data…
Axis ranges scale to the data — no model is clamped to an edge. Pricing comes from the shared index built on official provider docs; self-hosted / free models are staggered just off the $0 axis so their labels don't pile up, and they stay out of the frontier math because there is no per-token price to compare.
Benchmark explorer
Drag the sliders to weight what matters to you. The ranking updates as you move them. Set an irrelevant dimension to zero. If you weight a dimension for which a model has no data, that model receives zero for the dimension rather than having the gap ignored.
Loading…
Also useful
Frequently asked questions
Which AI model has the best capability-per-dollar?
Read the frontier line, then the cost-per-point column. On the default capability axis DeepSeek V4-Flash currently buys the most capability per dollar of any model with a published per-token price, and DeepSeek V4-Pro is the next step up the line. Most priced models are not on the frontier at all: something cheaper already matches or beats them. The summary above the table counts them from the live data, so it stays right when prices move.
What is the value frontier?
Plain arithmetic, not an opinion. A model sits on the frontier when no other model with a published API price is both cheaper and at least as good on the metric you picked. Everything else is beaten by something cheaper, and the table names which model beats it. Models with no per-token API rate are left out of that math — there is no price to compare, which is not the same as being free to run.
What does the vertical axis measure?
Your choice. The default is benchr's 0–100 capability weighting: coding 40%, reasoning 40%, writing 20%, where coding and reasoning read sourced benchmark fields and the rest are editorial ratings. You can switch it to SWE-bench Verified or GPQA Diamond, which plot only the models whose provider actually published that figure — a missing number keeps a model off the chart instead of scoring it zero. The explorer below changes the weighting itself.
Why are Llama/Phi plotted at zero price?
The chart uses $0 when the model record has no per-token API rate. That means the chart has no API fee to plot; it does not mean the model is free to run. Hardware, hosting, and operations are outside this chart.
Can I share these charts?
Yes. Share the page URL, or use the public models.json file under CC BY 4.0. Third-party iframe embedding is blocked by the site's security headers.