Release Timeline · Reviewed August 21, 2026
AI model release timeline
Major frontier and open-weight releases, newest first, with dated pricing and the key change. Reviewed against provider sources on August 21, 2026.
2026
August 18, 2026
GLM-5.3 Hosted API
Z.AI · Analysis →glm-5.3 lists 1M context, 128K maximum output, always-on reasoning, and $1.40 input / $4.40 output / $0.26 cached input per 1M. Z.AI announced weights for later release after safety hardening; hosted access and weight availability remain separate states.August 13, 2026
Gemini 3.7 Flash Stable
Google · Analysis →gemini-3.7-flash lists 1,048,576 context, 65,536 maximum output, and introductory $0.75/$3.75 standard pricing through December 31, 2026. The scheduled January rate is $1.50/$7.50.August 12, 2026
Grok 4.6 Frontier
xAI · Analysis →grok-4.6 lists 500K context and $2 input / $6 output / $0.50 cached input per 1M. xAI describes no text-output limit but publishes no numeric cap, so BenchR leaves that numeric field blank.
July 21, 2026
Gemini 3.6 Flash Stable
Google · Coverage →
Stable Gemini API model. Google lists $1.50 input, $7.50 output, and $0.15 cached input per 1M tokens, with a 1,048,576-token input window and 65,536-token output limit.
July 21, 2026
Gemini 3.5 Flash-Lite Stable
Google
Stable high-volume automation model. Google lists $0.30 input, $2.50 output, and $0.03 cached input per 1M tokens.
July 16, 2026
Kimi K3 Open weights
Moonshot AI · Analysis →kimi-k3 lists 2.8T total / 104B active parameters, 896 experts, 1,048,576 context, and open weights under the Kimi K3 License. Hosted global API pricing is $3 cache-miss input, $0.30 cache-hit input, and $15 output per 1M.
July 9, 2026
Meta Muse Spark 1.1
Preview
Meta
Public preview through Meta Model API. Meta describes a 1M-token multimodal model for creative and long-context workflows. No official public API price was posted at launch, so benchr records the release in model-figures.json but does not rank it in the pricing tool yet.
July 8, 2026
Grok 4.5
Frontier
xAI · Pricing →
xAI API model
grok-4.5. $2 input, $6 output, and $0.50 cached input per 1M tokens; 500K context. xAI has not published benchmark tables or max-output limits for this model yet.
July 1, 2026
GPT-5.6 — Sol, Terra, Luna
Frontier
OpenAI · Coverage →
Generally available since July 9, 2026. OpenAI lists Sol, Terra, and Luna in the API changelog, models page, and pricing page. Sol: $5/$30 per 1M, 1.05M context, 128K output, API
gpt-5.6-sol. OpenAI cut Terra to $2/$12 and Luna to $0.20/$1.20 on July 30, 2026 (previously $2.50/$15 and $1/$6); all three keep the same 1.05M context and 128K max output.
July 1, 2026
Claude Sonnet 5
Frontier
Anthropic · Coverage →
$2/$10 per 1M — launch pricing Anthropic confirmed as standard on August 31, 2026, cancelling the planned $3/$15 rise (cached $0.20; batch at 50% off). 1M context, 128K output — double Sonnet 4.6's 64K. Safety-classified requests return an explicit refusal; another-model retry requires configured logic. Benchmark figures shown are provider-reported. API:
claude-sonnet-5.
July 1, 2026
Claude Fable 5 — restored
Anthropic · Coverage →
The US export-control suspension in place since June 12 was lifted for all customers; AWS restored Bedrock access the same day — the same day Anthropic launched Claude Sonnet 5. Pricing and specs unchanged: $10/$50 per 1M, 1M context, 128K output. Mythos 5 remains Project-Glasswing-only, unaffected by this restoration.
June 30, 2026
Gemini Omni Flash Preview & Gemini 3.1 Flash Lite Image
Google
Omni Flash Preview adds text/image/video/audio input plus short video output with a 1,048,576-token context window. Gemini 3.1 Flash Lite Image is a GA image model with 65,536-token context and 4,096-token max output. Both are tracked in model-figures.json; media pricing is kept outside the ranked chat/coding model table.
June 26, 2026 (preview)
GPT-5.6 — Sol, Terra, Luna
Frontier
OpenAI · Coverage →
Announced $5/$30 (Sol), $2.50/$15 (Terra), $1/$6 (Luna) per 1M during a limited preview to about 20 government-approved partners. The model family became generally available on July 9, 2026; see the July entry above for live API pricing and specs.
June 16, 2026
GLM-5.2
Z.AI · Pricing →
1M-token context, 128K max output. Standard API pricing is $1.40 input, $4.40 output, and $0.26 cache-hit input per 1M tokens. Added to the ranked pricing and comparison data.
June 9, 2026
Claude Fable 5 & Claude Mythos 5
Frontier
Anthropic · Coverage →
$10/$50 per 1M tokens. 1M context, 128K output. Safety-classified requests return an explicit refusal; another-model retry requires configured logic. Mythos 5 is Glasswing-only. Fable was suspended June 12 and restored worldwide July 1, 2026.
June 1, 2026
MiniMax M3
MiniMax · Pricing →
1M-context hosted text model. Standard tier is $0.30 input, $1.20 output, and $0.06 cache-hit input per 1M tokens, with long-context and priority tiers priced separately.
May 28, 2026
Claude Opus 4.8
Frontier
Anthropic · Review →
$5/$25 per 1M tokens standard; $10/$50 Fast Mode. SWE-Bench Pro 69.2% (up from 64.3% on 4.7). Major alignment and honesty improvements. API:
claude-opus-4-8.
May 19, 2026 (GA)
Gemini 3.5 Flash
Frontier
Google · Review →
$1.50/$9.00 per 1M; free API tier. 1,048,576-token context and 65,536-token output limit. Google does not publish one universal throughput figure; measure latency in the target region. Announced at Google I/O 2026.
April 2026
Claude Opus 4.7
Frontier
Anthropic · Review →
$5/$25 per 1M. SWE-Bench Pro 64.3%. Now superseded by 4.8 but still fully supported.
April 24, 2026
GPT-5.5
Frontier
OpenAI · Review →
$5/$30 per 1M standard; $2.50/$15 batch. 1.05M context, 128K output. Built for agentic coding and computer use. Successor to GPT-5.
April 24, 2026
DeepSeek V4-Pro & V4-Flash
Open (MIT)
DeepSeek · Review →
Launch rates: V4-Pro $0.435/$0.87 per 1M, cache hit $0.003625; V4-Flash $0.14/$0.28, the cheapest commercial API at the time. 1M context, 384K output. MIT license. Self-hosted or API. DeepSeek repriced the line on August 16, 2026 — see the price history.
April 22, 2026
Qwen 3.6
Open (Apache 2.0)
Alibaba · Review →
Apache 2.0. Dense 27B and MoE 35B-A3B variants. 262K native context (~1M extended). Strong multilingual. Self-hosted — no verified per-token API price.
April 20, 2026
Kimi K2.6
Open (Modified MIT)
Moonshot AI · Review →
$0.95/$4.00 per 1M; cached $0.16. 256K context. 1T total / 32B active MoE. Good multilingual performance.
April 2026
Grok 4.3
Frontier
xAI · Review →
$1.25/$2.50 per 1M; cached $0.20. 1M context. Real-time data via X. Very cheap output tokens for agentic loops.
April 2026
Mistral Medium 3.5
Open (Modified MIT)
Mistral
$1.50/$7.50 per 1M. Public preview. 256K context. Multimodal + reasoning. Self-host on 4 GPUs.
March 10, 2026
Grok 4.20
Frontier
xAI
$2/$6 per 1M. 2M context — the largest of any flagship at release. Public beta from February 17; exited beta mid-March with Auto/Fast/Expert/Heavy modes. Superseded by Grok 4.3 in April.
March 5, 2026
GPT-5.4
Frontier
OpenAI
$2.50/$15 per 1M; cached $0.25. Context up to 1M, 128K output. Built-in computer use (OSWorld-Verified 75%); tuned for finance work and launched alongside ChatGPT for Excel. Thinking + Pro first, mini and nano on March 17. Superseded as flagship by GPT-5.5.
March 3, 2026
GPT-5.3 Instant
OpenAI
ChatGPT default model from March 3 to May 5, 2026, when GPT-5.5 Instant replaced it. Paid users keep it roughly three months from that date.
February 2026
Gemini 3.1 Pro
Frontier
Google · Review →
$2/$12 per 1M (above 200K: $4/$18). 1M context, 64K output. Deepest reasoning in the Gemini family. Replaced Gemini 3 Pro.
February 16, 2026
Qwen 3.5
Open
Alibaba
Flagship 397B MoE, 17B active; family of nine sizes, most Apache 2.0. Native 262K context. Superseded by Qwen3.6 in April.
February 5, 2026
Claude Opus 4.6
Frontier
Anthropic
$5/$25 per 1M. First Opus with the 1M-token context window (beta at launch), 128K output. Superseded by Opus 4.7 in April, then 4.8 in May; still available as a legacy model.
2025
December 2, 2025
Mistral Large 3
Open (Apache 2.0)
Mistral · Review →
$0.50/$1.50 per 1M. Apache 2.0 — true open commercial license. 675B total / 41B active MoE. 256K context. API id: mistral-large-2512.
November 2025
Claude Haiku 4.5
Small
Anthropic · Review →
$1/$5 per 1M. 200K context and 64K max output. High-volume routing and classification candidate; latency requires a deployment test.
September 2025
Claude Sonnet 4.6
Mid
Anthropic · Review →
$3/$15 per 1M. 200K context. 95 tok/s. The production daily-driver between Haiku and Opus.
August 7, 2025
GPT-5
Frontier
OpenAI · Review →
$1.25/$10 per 1M. 1.05M context, 128K output. The everyday OpenAI production pick at a rational price point.
April 5, 2025
Llama 4 Scout & Maverick
Open (Community)
Meta · Review →
Free under Meta's Community License. Scout: 10M token context, 17B/109B; Maverick: 1M context, 17B/400B, 128 experts. Self-hosted at no API cost.
January 2025
Phi-4
Small
Microsoft
MIT license. 14B parameter dense model. 16K context. Built for edge/local inference. Punches above its weight on MMLU and math.
Frequently asked questions
What is the newest AI model in 2026?
As of August 21, 2026, the newest confirmed model releases are GLM-5.3 on August 18, Gemini 3.7 Flash on August 13, and Grok 4.6 on August 12. Each has a source-linked model record and bilingual analysis.
Which AI models were released in 2026?
In 2026 through August 21: GLM-5.3, Gemini 3.7 Flash, and Grok 4.6 lead the August entries; Kimi K3, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, GPT-5.6 general availability, and the earlier releases follow in reverse chronology above.