The major labs ship something significant every few weeks. This page lists significant releases and material platform changes since early 2026, in reverse chronological order. Every August entry below was rechecked against an official provider source on August 21. For deeper coverage of each model, see the linked article or the comparison tool.
August 2026
- OpenAI Model Spec update — On August 18, OpenAI updated the Model Spec. This is a behavior-policy change, not a new production model, so BenchR records it here without inventing an API SKU. Official source.
- GLM-5.3 — Z.AI released the hosted
glm-5.3endpoint on August 14 with 1M context, 128K maximum output, always-on reasoning, and $1.40/$4.40 list pricing. The provider announced that weights would follow after safety hardening; they are not recorded as available yet. BenchR analysis · Official source. - Grok 4.6 in GitHub Copilot — xAI announced Grok 4.6 availability in GitHub Copilot on August 14, two days after the API launch. Official source.
- Gemini 3.7 Flash — Google made
gemini-3.7-flashgenerally available on August 13 with 1,048,576 input tokens, 65,536 maximum output, and introductory $0.75/$3.75 pricing through December 31. BenchR analysis · Official source. - Grok 4.6 — xAI launched
grok-4.6on August 12 with a 500K context window, text and image input, tool use, and base pricing of $2 input, $6 output, and $0.50 cached input per million tokens. BenchR analysis · Official source.
July 2026
- Kimi K3 — Moonshot AI released Kimi K3 on July 16 and later published the weights. The official record lists 2.8T total and 104B active parameters, 1,048,576 context tokens, and hosted pricing of $3 cache-miss input, $0.30 cache-hit input, and $15 output per million. BenchR analysis · Official source.
- Claude Opus 5 — Anthropic released Claude Opus 5 on July 24, 2026. The article covers its verified API record and migration implications. benchr analysis · Official source.
- GPT-Realtime-2.1 — OpenAI announced GPT-Realtime-2.1 on July 6, 2026 for low-latency voice and multimodal experiences. benchr analysis · Official source.
- GPT-Realtime-2.1 mini — OpenAI announced GPT-Realtime-2.1 mini on July 6, 2026 as its lower-cost realtime companion. benchr analysis · Official source.
- GPT-Live-1 — OpenAI introduced GPT-Live-1 on July 8, 2026 for ChatGPT Voice and said API availability was planned. benchr analysis · Official source.
- GPT-Live-1 mini — OpenAI introduced GPT-Live-1 mini on July 8, 2026 as part of the ChatGPT Voice rollout. benchr analysis · Official source.
- Gemini 3.5 Flash-Lite — Google released Gemini 3.5 Flash-Lite as generally available on July 21, 2026. benchr analysis · Official source.
- Leanstral 1.5 — Mistral released Leanstral 1.5 on July 2, 2026 as an Apache-2.0 formal-verification specialist. benchr analysis · Official source.
- Qwen-Audio 3.0 TTS Plus — Alibaba Cloud released Qwen-Audio 3.0 TTS Plus on July 14, 2026. benchr analysis · Official source.
- Qwen-Audio 3.0 TTS Flash — Alibaba Cloud released Qwen-Audio 3.0 TTS Flash on July 14, 2026. benchr analysis · Official source.
- Muse Image — Meta launched Muse Image on July 7, 2026 and made it available in Meta AI. benchr analysis · Official source.
- Muse Video Preview — Meta previewed Muse Video on July 7, 2026 alongside Muse Image. benchr analysis · Official source.
- Gemini 3.6 Flash — Released July 21, 2026 as a stable Gemini API model. Google lists
gemini-3.6-flashat $1.50 input, $7.50 output, and $0.15 cached input per 1M tokens, with a 1,048,576-token input window and 65,536-token output limit. Source: Google. benchr analysis. - Meta Muse Spark 1.1 — Released July 9, 2026 in public preview through Meta Model API. Meta positions Spark as a 1M-token multimodal model for long-context creative workflows; no public API price was posted in the announcement, so benchr tracks it in model-figures.json but keeps it out of the ranked pricing tool for now. Source: Meta.
- Grok 4.5 — Released to the xAI API July 8, 2026. xAI lists
grok-4.5with a 500K context window at $2 input, $6 output, and $0.50 cached input per 1M tokens. xAI does not publish official benchmark tables or a max-output limit for this model yet. benchr review · Source: xAI · benchr pricing. - GPT-5.6 — generally available — July 9, 2026. OpenAI's API changelog now lists GPT-5.6 Sol, Terra, and Luna as released to the API. Official prices are unchanged from the June preview: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M; OpenAI's models page lists all three with 1.05M context and 128K max output. Source: OpenAI. benchr writeup · pricing.
- Claude Sonnet 5 — Released July 1, 2026. Anthropic's second Mythos-class-architecture model after Claude Fable 5, priced at an introductory $2/$10 per million tokens through August 31, then $3/$15 from September 1, with a 1M context window and a 128K max output — double Sonnet 4.6's 64K. Anthropic reports SWE-bench Verified 89.4% and GPQA Diamond 92.0%. Safety-classified requests return an explicit refusal; another-model retry requires configured application or product logic. Source: Anthropic. benchr coverage · pricing.
- Claude Fable 5 restored to all customers — July 1, 2026. The U.S. Commerce Department export-control review that suspended Fable 5 on June 12 concluded, and Anthropic restored access for every customer, with AWS restoring Bedrock access the same day; Claude Mythos 5 remains limited to Project Glasswing. Landed the same day as the Claude Sonnet 5 launch above. Source: Anthropic. benchr coverage.
June 2026
- Gemini Omni Flash Preview — Released June 30, 2026. Google's new preview model accepts text, image, video, and audio input and can produce text plus short video clips. The Gemini docs list a 1,048,576-token context window and pricing of $1.50 input, $9 text output, and $17.50 video output per 1M tokens. Source: Google.
- Gemini 3.1 Flash Lite Image — Released June 30, 2026. Google lists this Nano Banana 2 Lite image model as generally available with a 65,536-token context window, 4,096-token max output, $0.25 input, $1.50 text output, and $30 image output per 1M tokens. Source: Google.
- GLM-5.2 — Released June 16, 2026 by Z.AI. The API model
glm-5.2lists a 1M-token context window, 128K max output, and pricing at $1.40 input, $4.40 output, and $0.26 cache-hit input per 1M tokens. benchr review · Source: Z.AI · benchr pricing. - MiniMax M3 — Released June 1, 2026. MiniMax lists
MiniMax-M3as a 1M-context text model with standard pricing at $0.30 input, $1.20 output, and $0.06 cache-hit input per 1M tokens; long-context and priority tiers cost more. Source: MiniMax. benchr pricing. - GPT-5.6 — Sol, Terra, and Luna — Previewed June 26, 2026, then generally available July 9. The preview introduced Sol, Terra, Luna, and the max/ultra reasoning modes under government-gated partner access; the July 9 API changelog moved the family into general availability. Source: OpenAI.
- Mistral OCR 4 — Released June 23, 2026. State-of-the-art document OCR: paragraph-level bounding boxes, typed-block labels, 170 languages. Priced per page ($4 / 1,000 pages standard, $2 batch, $5 Document AI), not per token. Source: Mistral.
- Qwen-AgentWorld (35B-A3B and 397B-A17B) — Released June 24, 2026. Open-weight (Apache-2.0) “language world models” for agent-environment simulation across seven domains; 256K context, download-only. Source: Qwen.
- Kimi K2.7-Code — Moonshot's coding-specialist successor to K2.6, built on the same model. $0.95/1M input (cache miss), $0.19 cache hit, $4/1M output, 256K context, Modified MIT; a Highspeed tier doubles the rates ($1.90/$8) for ~180 tok/s. Moonshot publishes no official release date. Source: Moonshot. benchr pricing.
- Claude Fable 5 and Claude Mythos 5 — Released June 9, 2026. Fable 5 lists at $10/$50 per million tokens with 1M context and 128K max output; safety-classified requests return an explicit refusal, and a retry on another model must be configured. Mythos 5 stays restricted to Project Glasswing. Source: Anthropic. benchr coverage. Timeline: both models were suspended June 12, controls were lifted June 30, and access returned July 1. Fable 5 is generally available worldwide; Mythos 5 remains limited to approved Glasswing organizations.
May 2026
- Grok Build 0.1 — Reached the xAI API in public beta around May 29, 2026. Agentic coding model, 256K context, $1.00/$2.00 per 1M. Source: xAI.
- Qwen3.7-Max — Announced May 21, 2026 and now available through Alibaba Cloud Model Studio. The announcement still described API access as “coming soon,” so benchr does not assign an unsupported first-availability date. Singapore/international list pricing is $2.50/$7.50 per 1M tokens with a 1M input tier. Announcement; current availability; pricing.
- Claude Opus 4.8 — Released May 28, 2026. Anthropic's newest flagship. Same $5/$25 pricing as 4.7, 69.2% on SWE-Bench Pro (up from 64.3%), with a notable honesty and alignment gain. benchr review.
- Gemini 3.5 Flash — Released May 19, 2026. Beats Gemini 3.1 Pro on most coding and agentic benchmarks. $1.50 input / $9 output per million tokens. 1M context. Source: Google. benchr review.
April 2026
- Gemini 3.1 Flash TTS Preview — Google launched Gemini 3.1 Flash TTS Preview on April 15, 2026. benchr analysis · Official source.
- Gemma 4 26B A4B IT — Google released gemma-4-26b-a4b-it on April 2, 2026 as part of the Gemma 4 launch. benchr analysis · Official source.
- Gemma 4 31B IT — Google released gemma-4-31b-it on April 2, 2026 as part of the Gemma 4 launch. benchr analysis · Official source.
- Mistral Medium 3.5 — Released April 28, 2026. 128B dense, 256K context, vision built-in. Replaces Medium 3.1 + Magistral + Devstral 2 in one model. $1.50 / $7.50 per million tokens. Modified MIT license. Source: Mistral. benchr review.
- DeepSeek V4-Pro and V4-Flash — Released April 24, 2026. V4-Pro: 1.6T MoE, 49B active, 80.6% SWE-bench Verified. V4-Flash: 284B MoE, 13B active, $0.14 input. Source: DeepSeek. benchr review.
- GPT-5.5 — Released April 23, 2026. OpenAI's new flagship. $5 input / $30 output. 1M context. Strong on Terminal-Bench and FrontierMath. Source: OpenAI. benchr review.
- Claude Opus 4.7 — Released April 16, 2026. Same $5/$25 pricing as Opus 4.6, 87.6% SWE-bench Verified. New tokenizer produces up to 35% more tokens per request. Source: Anthropic. benchr review.
March 2026
- Mistral Small 4 — Released March 16, 2026. Open-weight (Apache-2.0) hybrid MoE, 119B total / 6B active, 256K context, text + image input. $0.15/$0.60 per 1M. Source: Mistral.
- GPT-5.4 family — Released March 5, 2026 (Thinking + Pro). GPT-5.4 mini and nano followed on March 17. $2.50/$15 for the main model, context up to 1M tokens, built-in computer use. Launched alongside ChatGPT for Excel. Source: OpenAI.
- GPT-5.3 Instant — Released March 3, 2026 as the ChatGPT default model. Replaced by GPT-5.5 Instant on May 5; paid users keep it for roughly three months from that date. Source: OpenAI release notes.
- Gemini 3 Pro deprecated — March 9, 2026. Replaced by Gemini 3.1 Pro Preview as the Pro tier. Source: Google changelog.
February 2026
- Gemini 3.1 Pro Preview — Released February 19, 2026. $2 input / $12 output per million tokens. Still the Pro tier; Gemini 3.5 Pro remains unreleased as of this writing. benchr review.
- Qwen 3.5 — Released February 16, 2026. Flagship 397B MoE, 17B active; the family spans nine sizes, most under Apache 2.0. Superseded by Qwen3.6 in April. Source: Qwen GitHub. benchr review.
What to watch
Claude Fable 5's return — resolved. The export-control review Commerce opened in June concluded on July 1, 2026, and Anthropic restored Fable 5 to all customers, with AWS restoring Bedrock access the same day. Claude Mythos 5 remains limited to Project Glasswing.
Next up. OpenAI has disclosed an upcoming model named Astra in a safety-planning announcement, but has not released a production model card, API identifier, or price. It remains a watch item rather than a BenchR model record. Z.AI's GLM-5.3 weights also remain behind the announced safety-hardening period, and Moonshot's Kimi K2.5/Moonshot V1 platform sunset is scheduled for August 31.
The open-weight tier (DeepSeek V4, Qwen 3.5/3.6, Mistral Medium 3.5) has closed the gap with closed labs to single-digit benchmark points on most evaluations. Coding and reasoning leaders are now distributed across both camps.
Changelog
- — Added the August official-source sweep, the July Kimi K3 record, dated lifecycle changes, and links to the four bilingual analyses.