AI model rankings
Compare 33 frontier, mid-tier, and open-weight models. Start with capability, or switch to Value to include price. Capability profiles are editorial judgments, not independent measurements. Speed stays blank unless a reproducible or clearly attributed figure exists.
| # | Model | benchr Rating | SWE-bench % | Input $/1M | Output $/1M | Context | Tok/s | Released |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropicfrontier | 9.8 | — | $10.00 | $50.00 | 1M | — | Jun 2026 |
| 2 | Claude Opus 5Anthropicfrontier | 9.6 | — | $5.00 | $25.00 | 1M | — | Jul 2026 |
| 3 | Claude Sonnet 5Anthropicfrontier | 9.5 | 89.4% | $2.00 | $10.00 | 1M | — | Jul 2026 |
| 4 | Claude Opus 4.8Anthropicfrontier | 9.5 | 88.6% | $5.00 | $25.00 | 1M | — | May 2026 |
| 5 | Claude Opus 4.7Anthropicfrontier | 9.5 | 87.6% | $5.00 | $25.00 | 1M | — | Apr 2026 |
| 6 | GPT-5.6OpenAIfrontier | 9.4 | 89.8% | $5.00 | $30.00 | 1.1M | — | Jul 2026 |
| 7 | GPT-5.5OpenAIfrontier | 9.3 | — | $5.00 | $30.00 | 1.1M | — | Apr 2026 |
| 8 | Grok 4.6xAIfrontier | 9.3 | — | $2.00 | $6.00 | 500K | — | Aug 2026 |
| 9 | Gemini 3.7 FlashGooglefrontier | 9.2 | — | $0.750 | $3.75 | 1.0M | — | Aug 2026 |
| 10 | Kimi K3Moonshot AIfrontier open | 9.2 | — | $3.00 | $15.00 | 1.0M | — | Jul 2026 |
| 11 | DeepSeek V4-ProDeepSeekfrontier open | 9.1 | 80.6% | $0.435 | $0.870 | 1M | — | Apr 2026 |
| 12 | Grok 4.5xAIfrontier | 9.1 | — | $2.00 | $6.00 | 500K | — | Jul 2026 |
| 13 | GLM-5.3Z.AIfrontier open | 9.1 | — | $1.40 | $4.40 | 1M | — | Aug 2026 |
| 14 | GPT-5.4OpenAIfrontier | 9.0 | — | $2.50 | $15.00 | 1.1M | — | Mar 2026 |
| 15 | Claude Sonnet 4.6Anthropicmid | 8.8 | 79.6% | $3.00 | $15.00 | 1M | — | Feb 2026 |
| 16 | GPT-5OpenAIfrontier | 8.8 | 74.9% | $1.25 | $10.00 | 400K | — | Aug 2025 |
| 17 | GLM-5.2Z.AIfrontier open | 8.8 | — | $1.40 | $4.40 | 1M | — | Jun 2026 |
| 18 | Gemini 3.6 FlashGooglemid | 8.7 | — | $1.50 | $7.50 | 1.0M | — | Jul 2026 |
| 19 | MiniMax M3MiniMaxfrontier open | 8.7 | — | $0.300 | $1.20 | 1M | — | Jun 2026 |
| 20 | Gemini 3.5 FlashGooglemid | 8.6 | — | $1.50 | $9.00 | 1.0M | — | May 2026 |
| 21 | Gemini 3.1 ProGooglefrontier | 8.6 | 80.6% | $2.00 | $12.00 | 1M | — | Feb 2026 |
| 22 | Mistral Medium 3.5Mistralfrontier open | 8.6 | — | $1.50 | $7.50 | 256K | — | Apr 2026 |
| 23 | Qwen3.6-27BAlibabaopen | 8.6 | 77.2% | Free | Free | 262K | — | Apr 2026 |
| 24 | DeepSeek V4-FlashDeepSeekopen | 8.5 | 79.0% | $0.140 | $0.280 | 1M | — | Apr 2026 |
| 25 | Gemini 3.5 Flash-LiteGooglefrontier | 8.2 | — | $0.300 | $2.50 | 1.0M | — | Jul 2026 |
| 26 | Grok 4.3xAIfrontier | 8.2 | — | $1.25 | $2.50 | 1M | — | — |
| 27 | Llama 4 MaverickMetafrontier open | 8.0 | — | Free | Free | 1M | — | Apr 2025 |
| 28 | Kimi K2.6Moonshot AIfrontier open | 7.8 | 80.2% | $0.950 | $4.00 | 262K | — | — |
| 29 | Mistral Large 3Mistralfrontier open | 7.8 | — | $0.500 | $1.50 | 256K | — | Dec 2025 |
| 30 | Claude Haiku 4.5Anthropicsmall | 7.6 | 73.3% | $1.00 | $5.00 | 200K | — | Oct 2025 |
| 31 | Phi-4Microsoftsmall open | 7.4 | — | Free | Free | 16K | — | Dec 2024 |
| 32 | GPT-5 MiniOpenAIsmall | 7.3 | — | $0.250 | $2.00 | 400K | — | Aug 2025 |
| 33 | Llama 4 ScoutMetaopen | 7.3 | — | Free | Free | 10M | — | Apr 2025 |
How the benchr Rating works
The Rank by toggle gives the benchr Rating two meanings. Quality uses capability alone. Value also includes listed API rates. Neither is a lab measurement or a reader poll; both are transparent editorial calculations, and no provider pays for placement.
Quality — the default
This view combines coding, reasoning, and writing. Its inputs come from the public record and are inspectable in models.json.
Value — the optional lens
This view blends capability with API-rate efficiency. Price is the average of listed input and output rates per million tokens. A self-hosted model with no per-token API rate receives the maximum API-price score. Hardware, operations, and hosted inference are excluded, so Value is not a total-cost comparison.
Both run in assets/js/models.js, so you can read and verify them yourself. Each produces a 0–100 value shown on a 0–10 scale, and price is shown directly in the Input/Output columns regardless of mode. For verified official pricing and benchmark figures, see model-figures.json.
Methodology as of June 1, 2026. The formula may change as the model field changes; check the changelog for updates.
Frequently asked questions
What is the benchr Rating?
By default it's pure capability (coding, reasoning, writing). A “Rank by” toggle switches it to a Value lens that also folds in price efficiency (most capability per dollar). Both run in open JavaScript you can read in assets/js/models.js; price is shown in its own columns too.
Which AI model is best in 2026?
It depends on the task and budget. The default benchr Rating is an editorial capability lens, while Value blends that lens with listed prices. Use the rankings to form a shortlist and validate it on representative tasks; no score on this page establishes a universal winner.
Are rankings ever paid or sponsored?
No. Rankings are computed purely from data in models.json. No provider has paid for placement. See editorial standards for the full policy.
How often is the data updated?
The updated field in models.json shows the last data refresh. Major releases and price changes are added after they are checked against provider documentation. Spot an error? File a correction.
What is the difference between a benchmark result and a capability profile?
A published benchmark result names the evaluation variant and its source. The 0–100 capability profiles are benchr editorial judgments informed by cited public evidence; they are not benchmark results measured or independently reproduced by benchr. Missing official values stay blank. The methodology page explains the distinction, and provider facts are tracked in model-figures.json.