Leaderboard · Updated June 2026

Cheapest AI model API leaderboard

Every model in benchr's index ranked by blended token cost — cheapest first. Open-weight models with no per-token fee are listed at the top. Data from official provider docs, updated June 2026.

Data from models.json Neutral reference — never sponsored
Rank Model Provider Blended $/1M Input $/1M Output $/1M
#1Qwen3.6-27BAlibabaSelf-hosted (free license)Self-hostedSelf-hosted
#2Llama 4 MaverickMetaSelf-hosted (free license)Self-hostedSelf-hosted
#3Llama 4 ScoutMetaSelf-hosted (free license)Self-hostedSelf-hosted
#4Phi-4MicrosoftSelf-hosted (free license)Self-hostedSelf-hosted
#5DeepSeek V4-FlashDeepSeek$0.210$0.140$0.280
#6DeepSeek V4-ProDeepSeek$0.652$0.435$0.870
#7Mistral Large 3Mistral$1.00$0.500$1.50
#8GPT-5 MiniOpenAI$1.12$0.250$2.00
#9Grok 4.3xAI$1.88$1.25$2.50
#10Kimi K2.6Moonshot AI$2.48$0.950$4.00
#11Claude Haiku 4.5Anthropic$3.00$1.00$5.00
#12Mistral Medium 3.5Mistral$4.50$1.50$7.50
#13Gemini 3.5 FlashGoogle$5.25$1.50$9.00
#14GPT-5OpenAI$5.62$1.25$10.00
#15Gemini 3.1 ProGoogle$7.00$2.00$12.00
#16Claude Sonnet 4.6Anthropic$9.00$3.00$15.00
#17Claude Opus 4.8Anthropic$15.00$5.00$25.00
#18Claude Opus 4.7Anthropic$15.00$5.00$25.00
#19GPT-5.5OpenAI$17.50$5.00$30.00

How the budget tier works in 2026

API pricing dropped faster than most teams expected. DeepSeek V4-Flash changed the market when it launched at $0.14/1M input — about 55% cheaper than GPT-5 Mini, the previous budget leader. Most companies haven't adjusted their cost models to account for just how cheap the bottom of the market has become.

The models in the sub-$1/1M input range aren't toy models. DeepSeek V4-Flash scores 79.0% on SWE-bench. GPT-5 Mini at $0.25 scores 48.0%. The correlation between price and capability has broken down at the lower tiers in ways that weren't true two years ago.

Self-hosted vs managed API: the real trade-off

Open-weight models from Meta, Alibaba, and Microsoft appear at $0 in this table because their licensing fee is zero. That doesn't mean they're free to run. Llama 4 Scout — a 109-billion-parameter model — requires multiple high-memory GPUs for production use. At low call volume, paying $0.14/1M on DeepSeek's API is cheaper than keeping a GPU warm for your own instance.

The crossover point is roughly 200–400M tokens per day depending on the model and your infrastructure efficiency. Below that threshold, managed APIs win on total cost even when DeepSeek or GPT-5 Mini is pricing them at fractions of a cent per thousand tokens. Above it, self-hosting becomes worth evaluating seriously.

Methodology

Blended price = (input price + output price) / 2. Self-hosted models show a $0 blended price reflecting zero licensing cost; actual infrastructure cost is excluded. All prices sourced from official provider documentation, verified June 3, 2026. Use the Cost Calculator to model your specific token mix.

When the cheapest model is not the cheapest choice

A low token price only wins when the model completes the task with a similar retry rate and similar human-review cost. If a cheaper model produces two extra failed generations for every successful answer, the apparent savings disappear quickly. Measure cost per accepted result, not cost per generated token.

For production routing, use this leaderboard as the first filter, then run a narrow evaluation on your own prompts. The practical test is simple: compare one budget model, one mid-tier default, and one frontier model on the same 100 to 300 examples, then price only the outputs your team would actually accept.

Update cadence and verification

Prices in this table are reviewed against provider documentation when a major model ships, when a provider announces a pricing change, or when benchr updates the shared model index. Because API prices can change quietly, the table should be treated as a decision aid, not a procurement contract. Before signing a long-term vendor agreement, re-check the provider's live pricing page and any enterprise discounts available to your account.

The table also separates token price from platform fit. A model can be cheapest and still be wrong for a team that needs existing SDK support, SOC review, cloud-region guarantees, or internal approval for a specific vendor. Those constraints belong in your evaluation sheet next to the token cost.

Frequently asked questions

What is the cheapest hosted LLM API in 2026?

DeepSeek V4-Flash at $0.14/1M input tokens and $0.28/1M output tokens. It's an open-weight MIT model also available for self-hosting. Among models you can only access via managed API, it holds the lowest per-token price of any frontier-adjacent model.

Are self-hosted open-weight models free?

The license fee is zero, but compute costs are not. Running Llama 4 Scout or Phi-4 requires renting GPU instances. At low call volume, hosted APIs are often cheaper. At high sustained volume — hundreds of millions of tokens daily — self-hosting typically wins on total cost.

How is blended price calculated?

Blended price is (input price + output price) / 2, assuming an equal mix of input and output tokens. Most real workloads are input-heavier, so treat blended price as a conservative estimate rather than your actual bill.