Pricing Index · Updated July 2026

AI model API pricing

Input, output, and caching rates for every major AI API — sourced from official provider docs and updated whenever prices change. 26 ranked chat/coding models from OpenAI, Anthropic, Google, xAI, Z.AI, MiniMax, DeepSeek, Mistral, Meta, and more, plus media-model records in the open data.

Data from models.json · verified July 13, 2026 Neutral reference — never sponsored

All models by price

Model Provider Input / 1M Output / 1M Cached input / 1M Details
Qwen3.6-27BAlibabaSelf-hostedSelf-hostedFull guide →
Llama 4 MaverickMetaSelf-hostedSelf-hostedFull guide →
Llama 4 ScoutMetaSelf-hostedSelf-hostedFull guide →
Phi-4MicrosoftSelf-hostedSelf-hostedFull guide →
DeepSeek V4-FlashDeepSeek$0.140$0.280$0.0028Full guide →
GPT-5 MiniOpenAI$0.250$2.00$0.025Full guide →
MiniMax M3MiniMax$0.300$1.20$0.060Full guide →
DeepSeek V4-ProDeepSeek$0.435$0.870$0.0036Full guide →
Mistral Large 3Mistral$0.500$1.50Full guide →
Kimi K2.6Moonshot AI$0.950$4.00$0.160Full guide →
Claude Haiku 4.5Anthropic$1.00$5.00$0.100Full guide →
GPT-5OpenAI$1.25$10.00Full guide →
Grok 4.3xAI$1.25$2.50$0.200Full guide →
GLM-5.2Z.AI$1.40$4.40$0.260Full guide →
Gemini 3.5 FlashGoogle$1.50$9.00$0.150Full guide →
Mistral Medium 3.5Mistral$1.50$7.50Full guide →
Gemini 3.1 ProGoogle$2.00$12.00$0.200Full guide →
Grok 4.5xAI$2.00$6.00$0.500Full guide →
GPT-5.4OpenAI$2.50$15.00$0.250Full guide →
Claude Sonnet 4.6Anthropic$3.00$15.00$0.300Full guide →
Claude Opus 4.8Anthropic$5.00$25.00$0.500Full guide →
Claude Opus 4.7Anthropic$5.00$25.00$0.500Full guide →
GPT-5.5OpenAI$5.00$30.00$0.500Full guide →
Claude Fable 5Anthropic$10.00$50.00$1.00Full guide →

Individual model pricing guides

Detailed breakdown per model: token costs, context limits, caching structures, use-case recommendations, and cost scenarios.

Pricing articles and analysis

Deeper reading on cost strategy, provider comparisons, and how to reduce your API bill.

How to read these pricing tables

API providers bill in tokens — roughly 750 words per 1,000 tokens for English text. Prices are listed per million tokens because per-token rates are too small to read ($0.000001250 is clearer as $1.25/1M). Most providers charge separately for input (text you send) and output (text the model generates).

Output tokens typically cost 4–12× more than input tokens because they require a separate forward pass per token. This means the input-to-output ratio of your workload matters enormously. A summarization task (high input, low output) has a very different cost profile than a story-generation task (low input, high output).

Three cost strategies for 2026

Route by task complexity. Use cheap small models (DeepSeek V4-Flash at $0.14/1M, GPT-5 Mini at $0.25/1M) for classification, routing, and simple extraction. Reserve expensive flagships for tasks that actually require their capability. Most pipelines can route 70–80% of calls to cheaper models.

Maximize prompt caching. If your system prompt or document context repeats across calls, caching cuts that portion of your input cost by 90% (Anthropic, Google) or ~99% (DeepSeek). This single optimization often reduces monthly bills more than switching models entirely.

Use Batch API for offline workloads. OpenAI, Anthropic, and others offer 50% discounts for asynchronous batch jobs — dataset evaluation, classification runs, bulk translation. If your workload can tolerate a 24-hour return window, batch pricing cuts your bill in half.

Use the Cost Calculator to model your specific token mix, or the Cheapest API leaderboard to find the lowest-cost model that meets your quality bar.

Frequently asked questions

What is the cheapest AI API in 2026?

DeepSeek V4-Flash at $0.14/1M input tokens and $0.28/1M output tokens is the cheapest hosted API. Open-weight models like Llama 4 Scout and Phi-4 have no per-token licensing fee but require self-hosting infrastructure.

What does input vs output pricing mean?

API providers charge separately for input tokens (text you send to the model) and output tokens (text the model generates). Output tokens typically cost 4–10× more per token than input. Most real applications have 3–8× more input tokens than output.

What is prompt caching and how much does it save?

Prompt caching stores your static system prompt or document context so that repeated calls reuse it at a steep discount. Anthropic offers 90% off, Google 90% off, DeepSeek ~99% off. OpenAI doesn't publish a caching discount for GPT-5. For agents with large repeated system prompts, caching can reduce your actual bill by 60–80%.

Are the model pricing pages on benchr accurate?

Prices are sourced directly from official provider documentation and verified against model-figures.json, benchr's single source of truth. The verified date is shown on each page. Prices change without notice — re-verify before making budget decisions.