Release Timeline · Reviewed August 21, 2026

AI model release timeline

Major frontier and open-weight releases, newest first, with dated pricing and the key change. Reviewed against provider sources on August 21, 2026.

Verified against official announcements No paid placements — every entry from official source

2026

August 18, 2026
GLM-5.3 Hosted API
Z.AI · Analysis →
glm-5.3 lists 1M context, 128K maximum output, always-on reasoning, and $1.40 input / $4.40 output / $0.26 cached input per 1M. Z.AI announced weights for later release after safety hardening; hosted access and weight availability remain separate states.
August 13, 2026
Gemini 3.7 Flash Stable
Google · Analysis →
gemini-3.7-flash lists 1,048,576 context, 65,536 maximum output, and introductory $0.75/$3.75 standard pricing through December 31, 2026. The scheduled January rate is $1.50/$7.50.
August 12, 2026
Grok 4.6 Frontier
xAI · Analysis →
grok-4.6 lists 500K context and $2 input / $6 output / $0.50 cached input per 1M. xAI describes no text-output limit but publishes no numeric cap, so BenchR leaves that numeric field blank.
July 21, 2026
Gemini 3.6 Flash Stable
Google · Coverage →
Stable Gemini API model. Google lists $1.50 input, $7.50 output, and $0.15 cached input per 1M tokens, with a 1,048,576-token input window and 65,536-token output limit.
July 21, 2026
Gemini 3.5 Flash-Lite Stable
Google
Stable high-volume automation model. Google lists $0.30 input, $2.50 output, and $0.03 cached input per 1M tokens.
July 16, 2026
Kimi K3 Open weights
Moonshot AI · Analysis →
kimi-k3 lists 2.8T total / 104B active parameters, 896 experts, 1,048,576 context, and open weights under the Kimi K3 License. Hosted global API pricing is $3 cache-miss input, $0.30 cache-hit input, and $15 output per 1M.
July 9, 2026
Meta Muse Spark 1.1 Preview
Meta
Public preview through Meta Model API. Meta describes a 1M-token multimodal model for creative and long-context workflows. No official public API price was posted at launch, so benchr records the release in model-figures.json but does not rank it in the pricing tool yet.
July 8, 2026
Grok 4.5 Frontier
xAI · Pricing →
xAI API model grok-4.5. $2 input, $6 output, and $0.50 cached input per 1M tokens; 500K context. xAI has not published benchmark tables or max-output limits for this model yet.
July 1, 2026
GPT-5.6 — Sol, Terra, Luna Frontier
OpenAI · Coverage →
Generally available since July 9, 2026. OpenAI lists Sol, Terra, and Luna in the API changelog, models page, and pricing page. Sol: $5/$30 per 1M, 1.05M context, 128K output, API gpt-5.6-sol. OpenAI cut Terra to $2/$12 and Luna to $0.20/$1.20 on July 30, 2026 (previously $2.50/$15 and $1/$6); all three keep the same 1.05M context and 128K max output.
July 1, 2026
Claude Sonnet 5 Frontier
Anthropic · Coverage →
$2/$10 per 1M — launch pricing Anthropic confirmed as standard on August 31, 2026, cancelling the planned $3/$15 rise (cached $0.20; batch at 50% off). 1M context, 128K output — double Sonnet 4.6's 64K. Safety-classified requests return an explicit refusal; another-model retry requires configured logic. Benchmark figures shown are provider-reported. API: claude-sonnet-5.
July 1, 2026
Claude Fable 5 — restored
Anthropic · Coverage →
The US export-control suspension in place since June 12 was lifted for all customers; AWS restored Bedrock access the same day — the same day Anthropic launched Claude Sonnet 5. Pricing and specs unchanged: $10/$50 per 1M, 1M context, 128K output. Mythos 5 remains Project-Glasswing-only, unaffected by this restoration.
June 30, 2026
Gemini Omni Flash Preview & Gemini 3.1 Flash Lite Image
Google
Omni Flash Preview adds text/image/video/audio input plus short video output with a 1,048,576-token context window. Gemini 3.1 Flash Lite Image is a GA image model with 65,536-token context and 4,096-token max output. Both are tracked in model-figures.json; media pricing is kept outside the ranked chat/coding model table.
June 26, 2026 (preview)
GPT-5.6 — Sol, Terra, Luna Frontier
OpenAI · Coverage →
Announced $5/$30 (Sol), $2.50/$15 (Terra), $1/$6 (Luna) per 1M during a limited preview to about 20 government-approved partners. The model family became generally available on July 9, 2026; see the July entry above for live API pricing and specs.
June 16, 2026
GLM-5.2
Z.AI · Pricing →
1M-token context, 128K max output. Standard API pricing is $1.40 input, $4.40 output, and $0.26 cache-hit input per 1M tokens. Added to the ranked pricing and comparison data.
June 9, 2026
Claude Fable 5 & Claude Mythos 5 Frontier
Anthropic · Coverage →
$10/$50 per 1M tokens. 1M context, 128K output. Safety-classified requests return an explicit refusal; another-model retry requires configured logic. Mythos 5 is Glasswing-only. Fable was suspended June 12 and restored worldwide July 1, 2026.
June 1, 2026
MiniMax M3
MiniMax · Pricing →
1M-context hosted text model. Standard tier is $0.30 input, $1.20 output, and $0.06 cache-hit input per 1M tokens, with long-context and priority tiers priced separately.
May 28, 2026
Claude Opus 4.8 Frontier
Anthropic · Review →
$5/$25 per 1M tokens standard; $10/$50 Fast Mode. SWE-Bench Pro 69.2% (up from 64.3% on 4.7). Major alignment and honesty improvements. API: claude-opus-4-8.
May 19, 2026 (GA)
Gemini 3.5 Flash Frontier
Google · Review →
$1.50/$9.00 per 1M; free API tier. 1,048,576-token context and 65,536-token output limit. Google does not publish one universal throughput figure; measure latency in the target region. Announced at Google I/O 2026.
April 2026
Claude Opus 4.7 Frontier
Anthropic · Review →
$5/$25 per 1M. SWE-Bench Pro 64.3%. Now superseded by 4.8 but still fully supported.
April 24, 2026
GPT-5.5 Frontier
OpenAI · Review →
$5/$30 per 1M standard; $2.50/$15 batch. 1.05M context, 128K output. Built for agentic coding and computer use. Successor to GPT-5.
April 24, 2026
DeepSeek V4-Pro & V4-Flash Open (MIT)
DeepSeek · Review →
Launch rates: V4-Pro $0.435/$0.87 per 1M, cache hit $0.003625; V4-Flash $0.14/$0.28, the cheapest commercial API at the time. 1M context, 384K output. MIT license. Self-hosted or API. DeepSeek repriced the line on August 16, 2026 — see the price history.
April 22, 2026
Qwen 3.6 Open (Apache 2.0)
Alibaba · Review →
Apache 2.0. Dense 27B and MoE 35B-A3B variants. 262K native context (~1M extended). Strong multilingual. Self-hosted — no verified per-token API price.
April 20, 2026
Kimi K2.6 Open (Modified MIT)
Moonshot AI · Review →
$0.95/$4.00 per 1M; cached $0.16. 256K context. 1T total / 32B active MoE. Good multilingual performance.
April 2026
Grok 4.3 Frontier
xAI · Review →
$1.25/$2.50 per 1M; cached $0.20. 1M context. Real-time data via X. Very cheap output tokens for agentic loops.
April 2026
Mistral Medium 3.5 Open (Modified MIT)
Mistral
$1.50/$7.50 per 1M. Public preview. 256K context. Multimodal + reasoning. Self-host on 4 GPUs.
March 10, 2026
Grok 4.20 Frontier
xAI
$2/$6 per 1M. 2M context — the largest of any flagship at release. Public beta from February 17; exited beta mid-March with Auto/Fast/Expert/Heavy modes. Superseded by Grok 4.3 in April.
March 5, 2026
GPT-5.4 Frontier
OpenAI
$2.50/$15 per 1M; cached $0.25. Context up to 1M, 128K output. Built-in computer use (OSWorld-Verified 75%); tuned for finance work and launched alongside ChatGPT for Excel. Thinking + Pro first, mini and nano on March 17. Superseded as flagship by GPT-5.5.
March 3, 2026
GPT-5.3 Instant
OpenAI
ChatGPT default model from March 3 to May 5, 2026, when GPT-5.5 Instant replaced it. Paid users keep it roughly three months from that date.
February 2026
Gemini 3.1 Pro Frontier
Google · Review →
$2/$12 per 1M (above 200K: $4/$18). 1M context, 64K output. Deepest reasoning in the Gemini family. Replaced Gemini 3 Pro.
February 16, 2026
Qwen 3.5 Open
Alibaba
Flagship 397B MoE, 17B active; family of nine sizes, most Apache 2.0. Native 262K context. Superseded by Qwen3.6 in April.
February 5, 2026
Claude Opus 4.6 Frontier
Anthropic
$5/$25 per 1M. First Opus with the 1M-token context window (beta at launch), 128K output. Superseded by Opus 4.7 in April, then 4.8 in May; still available as a legacy model.

2025

December 2, 2025
Mistral Large 3 Open (Apache 2.0)
Mistral · Review →
$0.50/$1.50 per 1M. Apache 2.0 — true open commercial license. 675B total / 41B active MoE. 256K context. API id: mistral-large-2512.
November 2025
Claude Haiku 4.5 Small
Anthropic · Review →
$1/$5 per 1M. 200K context and 64K max output. High-volume routing and classification candidate; latency requires a deployment test.
September 2025
Claude Sonnet 4.6 Mid
Anthropic · Review →
$3/$15 per 1M. 200K context. 95 tok/s. The production daily-driver between Haiku and Opus.
August 7, 2025
GPT-5 Frontier
OpenAI · Review →
$1.25/$10 per 1M. 1.05M context, 128K output. The everyday OpenAI production pick at a rational price point.
April 5, 2025
Llama 4 Scout & Maverick Open (Community)
Meta · Review →
Free under Meta's Community License. Scout: 10M token context, 17B/109B; Maverick: 1M context, 17B/400B, 128 experts. Self-hosted at no API cost.
January 2025
Phi-4 Small
Microsoft
MIT license. 14B parameter dense model. 16K context. Built for edge/local inference. Punches above its weight on MMLU and math.
→ Deprecation tracker — which models are ending → Changelog — recent data updates → Full model rankings with benchr Rating → Recent releases — editorial coverage

Frequently asked questions

What is the newest AI model in 2026?

As of August 21, 2026, the newest confirmed model releases are GLM-5.3 on August 18, Gemini 3.7 Flash on August 13, and Grok 4.6 on August 12. Each has a source-linked model record and bilingual analysis.

Which AI models were released in 2026?

In 2026 through August 21: GLM-5.3, Gemini 3.7 Flash, and Grok 4.6 lead the August entries; Kimi K3, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, GPT-5.6 general availability, and the earlier releases follow in reverse chronology above.

benchr updates

Follow new-model coverage through the public release archive or RSS feed.