Leaderboard · Updated August 2026

Cheapest AI model API leaderboard

Hosted routes with both input and output rates, ranked by an equal-weight price index. Models with no per-token API rate are listed separately, not treated as $0 APIs.

Data from models.json No paid placement · commercial links are disclosed
Rank Model Provider Blended $/1M Input $/1M Output $/1M
Loading hosted API prices…

Models without a per-token API rate

These records may be available for self-hosting, but the shared index has no input/output API pair to rank. License terms and infrastructure cost are separate questions.

ModelProviderStatus in this rankingCost to evaluate
Loading unpriced model records…

How the budget tier works in 2026

The budget tier changed hands this month. DeepSeek moved its V4 models to peak and off-peak rates on August 16, 2026, and V4-Flash input went from a flat $0.14/1M to $0.44 during peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday) and $0.22 the rest of the week. On the blended index above, that hands first place to MiniMax M3 at $0.75 while DeepSeek V4-Flash sits at $0.88 — though at its off-peak rate the same model blends to $0.44 and is still the cheapest row in this set by a distance. The table shows the peak rate, because that is the rate you pay unless you deliberately schedule around it.

One newer option is not in this index yet: Z.ai published GLM-5.3-Flash on August 26, 2026 at $0.15/1M input and $0.50/1M output, halved to $0.075 and $0.25 through September 9. It is in benchr's verified ledger but has not been added to the tool set that feeds this ranking, so treat it as a lead to check rather than a row you can compare here.

The blended ranking also includes output price, so the order can differ from a list sorted on input alone.

Lower price does not imply that a model is limited to trivial work. DeepSeek reports 79.0% on SWE-bench for V4-Flash, while OpenAI reports 48.0% for GPT-5 Mini. Provider figures and token prices still need to be tested against the same workload before they are treated as a quality-per-dollar result.

Self-hosted vs managed API

An open-weight model can have no license fee and still be expensive to serve. Hardware, utilization, concurrency, context length, electricity or cloud rent, engineering, monitoring, and support all enter the total. Models without a per-token API pair stay outside the hosted ranking.

There is no universal traffic level at which self-hosting becomes cheaper. Get a hardware or cloud quote for the exact checkpoint and serving target, measure sustained throughput, and compare that total with the managed API bill for the same accepted workload.

Methodology

Blended price = (input price + output price) / 2. Only records with both rates are ranked. The equal-weight result is a comparison index, not a bill estimate. All listed prices are sourced from provider documentation; use the Cost Calculator to model your token mix.

When the cheapest model is not the cheapest choice

A low token price only wins when the model completes the task with a similar retry rate and similar human-review cost. If a cheaper model produces two extra failed generations for every successful answer, the apparent savings disappear quickly. Measure cost per accepted result, not cost per generated token.

For production routing, use this leaderboard as the first filter, then run a narrow evaluation on your own prompts. The practical test is simple: compare one budget model, one mid-tier default, and one frontier model on the same 100 to 300 examples, then price only the outputs your team would actually accept.

Update cadence and verification

Prices in this table are reviewed against provider documentation when a major model ships, when a provider announces a pricing change, or when benchr updates the shared model index. Because API prices can change quietly, the table should be treated as a decision aid, not a procurement contract. Before signing a long-term vendor agreement, re-check the provider's live pricing page and any enterprise discounts available to your account.

The table also separates token price from platform fit. A model can be cheapest and still be wrong for a team that needs existing SDK support, SOC review, cloud-region guarantees, or internal approval for a specific vendor. Those constraints belong in your evaluation sheet next to the token cost.

Frequently asked questions

What is the cheapest hosted LLM API in 2026?

It depends on the hour. In this dated benchr set, MiniMax M3 has the lowest blended rate at $0.30/1M input and $1.20/1M output. DeepSeek V4-Flash is cheaper still — $0.22/$0.66 — but only outside DeepSeek's peak window; inside it the same model costs $0.44/$1.32. Both are open-weight MIT models. Check the provider's live rate card before budgeting.

Are self-hosted open-weight models free?

No. A model may have no license fee, but serving it still requires hardware, electricity or cloud instances, engineering, monitoring, and capacity planning. The crossover with a managed API depends on the exact model, throughput, hardware, utilization, and operating costs; this page does not assume a universal volume threshold.

How is blended price calculated?

Blended price is (input price + output price) / 2, an equal-weight comparison index. It is not a forecast of your bill. Use the calculator with your own input/output ratio, cache rate, batch use, and retries.