Changelog · Data log

benchr changelog

Every update to benchr's data — new models added, prices verified, deprecations tracked, scoring methodology changes. Each entry is dated and linked to the official source. The raw version history lives in history.json.

Record reconciliation
The August 21 market pass and the August 24–28 pricing corrections merged into one record; Kimi K2.5 retirement confirmed

Two update lines ran in parallel in August and are now one record: the August 21 pass that added Gemini 3.7 Flash, Grok 4.6, GLM-5.3, and Kimi K3, and the August 24–28 passes that recorded the GPT-5.6 Sol and Gemini 3.6 Flash cuts and DeepSeek's August 16 move to peak/off-peak billing. Every disputed figure was re-read on the provider's page on August 31: Kimi K3's $3/$0.30/$15 rates and model card were confirmed, GLM-5.3's release date corrected to August 18, and Moonshot's docs now mark Kimi K2.5 and the Moonshot V1 series retired today with calls returning 404 — the tracker and the permanent record reflect that. Separately, Anthropic's pricing page now lists Claude Sonnet 5's $2/$10 launch rate as the standard price — the increase to $3/$15 scheduled for September 1 will not occur.

Price rise + new models
DeepSeek moved the V4 line to peak/off-peak billing, GLM-5.3-Flash and V4-Flash-Vision-Exp joined the ledger, the Assistants API flipped to retired

DeepSeek replaced flat pricing with peak and off-peak rates at 16:00 UTC on August 16, 2026, a change the August 24 pass did not catch. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; every other hour bills at half. V4-Flash went from a flat $0.14/$0.28 per 1M tokens to $0.44/$1.32 at peak and $0.22/$0.66 off-peak, and V4-Pro from $0.435/$0.87 to $1.32/$3.96 and $0.66/$1.98. Cache-hit input rose furthest: about twelvefold on V4-Pro at peak. Both moves are logged in history.json, and every dependent pricing page, leaderboard, comparison, and article was repriced in English and Arabic. Published an analysis of what the schedule costs in both languages.

Added three models the ledger did not carry: GLM-5.3 (Z.ai, August 18, $1.4/$4.4, same base model as GLM-5.2 with post-training gains), GLM-5.3-Flash (August 26, a 320B-A18B natively multimodal MoE under MIT, $0.15/$0.50 list with a launch promotion halving that through September 9, 2026), and DeepSeek-V4-Flash-Vision-Exp (August 21, experimental, priced like V4-Flash). All three sit in the verified ledger only; none has entered the tool layer that feeds the rankings, because a model this new has no benchr capability record.

Lifecycle record: OpenAI's Assistants API shut down on August 26, 2026 and now appears under Past deprecations on the official page, so the entry flips from deprecated to retired. The migration path OpenAI names is the Responses API plus the Conversations API. No other tracked shutdown date has passed; Google's gemini-robotics-er-1.6-preview is still due August 31.

Checked and unchanged: Anthropic, OpenAI, Google, xAI, and Mistral published no model or price change between August 24 and August 28, 2026.

Price changes + new models
GPT-5.6 Sol cut to $4/$20, Gemini 3.7 Flash and Grok 4.6 ledger records re-verified, two deadlines flipped to retired

OpenAI cut GPT-5.6 Sol on August 21 from $5/$30 to $4/$20 per 1M tokens, with cached input falling from $0.50 to $0.40. The changelog labels the rate promotional through at least November 21, 2026. Google's pricing table now shows Gemini 3.6 Flash at $0.75/$3.75 through December 31, 2026, half the $1.50/$7.50 recorded on July 22. Both changes are logged in history.json.

Added two models the ledger did not carry: Gemini 3.7 Flash, generally available August 13 at the same rate as 3.6 Flash and with no published benchmark table, and Grok 4.6, announced August 12 with a 500K context window at $2/$6 per 1M below 200K prompt tokens. Published an analysis of the Flash price parity and the January 1, 2027 return to full price, in English and Arabic.

Lifecycle record: OpenAI's gpt-5.2-chat-latest and gpt-5.3-chat-latest now sit in the provider's past-deprecations section, and Google's embedding-2-preview shutdown date has passed, so both flipped from deprecated to retired. Four records the dataset never carried were added — OpenAI's July 20 deprecation of nine legacy audio, realtime, and transcription IDs (removal January 20, 2027), Google's August 17 Imagen 4.0 shutdown, Google's August 31 shutdown of gemini-robotics-er-1.6-preview, and Anthropic's August 17 retirement of the experimental prompt tools API and Workbench. The tracker now also carries the OpenAI Assistants API deadline of August 26, 2026, which it had been missing.

Market update + full integration
Gemini 3.7 Flash, Grok 4.6, GLM-5.3, and Kimi K3 added across BenchR

A market-wide official-source pass added Gemini 3.7 Flash, Grok 4.6, GLM-5.3, and Kimi K3 to the verified ledger, selector and comparison data, bilingual analysis, pricing references, archive, release timeline, search, feeds, and sitemap. Prices, cached-input rates, context, output limits, API IDs, reasoning modes, modalities, tools, availability, and provider benchmark variants were rechecked from primary sources.

The update records scheduled price changes and incomplete availability states instead of flattening them: Gemini 3.7 Flash's introductory rates expire after December 31; Grok 4.6 has no numeric text-output cap published; GLM-5.3 weights were announced but remained behind a safety-hardening gate; and Kimi K3's hosted and open-weight deployment contracts are kept separate. OpenAI's upcoming Astra was not added because it was not released.

Deprecation update
Google marks the embedding preview shut down; Imagen 4 reaches its earliest date next

Google's current Gemini API deprecations table now renders embedding-2-preview as shut down, so benchr moved it to the permanent retired record. Added imagen-4.0-generate-001, imagen-4.0-ultra-generate-001, and imagen-4.0-fast-generate-001 with their August 17 earliest shutdown date and the provider-listed gemini-3.1-flash-image replacement.

OpenAI's August 10 listing for gpt-5.2-chat-latest and gpt-5.3-chat-latest remains deprecated rather than retired: its current deprecations page still shows the May 8 notice without explicit past-status confirmation.

Verified data + new guidance
DeepSeek alias migration guide, Gemini access error, and robotics preview record

Re-verified that DeepSeek's legacy deepseek-chat and deepseek-reasoner aliases were retired on July 24, then published a practical migration checklist. The lifecycle record, archive, tracker, and review now link to that guide.

Added Google Gemini permission_denied (403) to the error reference using Google's API-error documentation, and added gemini-robotics-er-2-preview to the factual model ledger. Google publishes its API ID and robotics description, but not the price, context limit, or release date used by benchr, so those fields remain blank.

Deprecation update
Gemini embedding preview added; August 10 deadlines re-verified

Added Google's embedding-2-preview lifecycle record to the open deprecations dataset. Google lists August 10, 2026 as the earliest shutdown date and gemini-embedding-2 as the replacement. Re-verified OpenAI's same-day removal schedule for gpt-5.2-chat-latest and gpt-5.3-chat-latest, whose listed replacement is gpt-5.6-sol.

Both August 10 records remain labelled deprecated for now. On the date itself, the official provider pages still present them as scheduled records rather than confirmed past retirements, so benchr does not claim runtime removal without that evidence.

Editorial + operating rules
OpenAI Presence buyer's guide and a stricter update/deploy workflow

Published an English and Arabic OpenAI Presence analysis based on OpenAI's July 22 product announcement, business-data commitments, and API data-control documentation. The article separates provider-reported outcomes from independent evidence, records the missing public price and self-serve path, and includes a procurement and pilot checklist.

Added a project-local Benchr operator skill and corrected stale instructions about article/model counts, Arabic social cards, AdSense loading, application schema, and the difference between a discovery-only request and an already-authorized update/build/deploy request.

New models
Market sweep adds current text, voice, image, video, open-weight, and specialist releases

Added Claude Opus 5 to model-figures.json and models.json after verifying its July 24 launch, claude-opus-5 API ID, $5/$25 per-1M pricing, 1M context window, and 128K max output against Anthropic's current documentation. Added the text-comparison record for Gemini 3.5 Flash-Lite ($0.30/$2.50, 1M context, 64K output).

The verified source-of-truth now also records GPT-Realtime-2.1 and mini, GPT-Live-1 and mini, Gemini 3.1 Flash TTS Preview, Gemma 4 26B A4B IT, Gemma 4 31B IT, Leanstral 1.5, Qwen-Audio 3.0 TTS Plus and Flash, Muse Image, and Muse Video Preview. Where a provider has not published a price, limit, benchmark, or API ID, the field remains null rather than guessed.

Deprecation update
July retirement records confirmed; later OpenAI chat snapshots separated

Re-verified the deprecation record against current OpenAI, Anthropic, and DeepSeek documentation. The July 23 OpenAI legacy IDs, DeepSeek's legacy aliases, and Claude Opus 4.7 fast mode are now recorded as retired. The August 10 shutdown for gpt-5.2-chat-latest and gpt-5.3-chat-latest is now a separate upcoming record. Claude Mythos Preview remains deprecated, but Anthropic no longer lists a retirement date, so benchr removed the prior June 30 date rather than calling it retired.

New model
Gemini 3.6 Flash added after Google's July 21 GA release

Added Gemini 3.6 Flash to model-figures.json and the model ranking layer after Google listed gemini-3.6-flash as a stable Gemini API model. Official figures verified July 22: $1.50 input / $7.50 output / $0.15 cached input per 1M on standard, $0.75/$3.75 batch and flex, 1,048,576 input tokens, and 65,536 output tokens. No official benchmark table was published, so hard benchmark fields remain null in the source-of-truth file. New coverage: Gemini 3.6 Flash launch analysis.

New models
Market-wide model gap fill: Grok 4.5, GLM-5.2, MiniMax M3, Meta Muse Spark 1.1, and new Gemini media models

Added six official-source model records to model-figures.json: Grok 4.5 from xAI, GLM-5.2 from Z.AI, MiniMax M3, Meta Muse Spark 1.1, Gemini Omni Flash Preview, and Gemini 3.1 Flash Lite Image. The ranked tool layer now includes the three text/API models with official token pricing: Grok 4.5, GLM-5.2, and MiniMax M3.

Meta's preview model and Google's media models remain in the source-of-truth figures file but are excluded from the chat/coding ranking because their pricing or modality does not map cleanly to that table. New pricing pages: Grok 4.5, GLM-5.2, and MiniMax M3.

Correction
Three factual corrections issued: Gemini 3.5 Pro retracted, GPT-5.6 GA claim walked back, Claude Sonnet 5 max output fixed

Gemini 3.5 Pro retracted. The June 30, 2026 claim that Google released Gemini 3.5 Pro was fabricated: both source URLs originally cited now 404, Google's Gemini API pricing and model docs still list Gemini 3.1 Pro as the current flagship, and Google DeepMind's own model hub lists "Gemini 3.5 Pro" as coming soon, not shipped. The review and pricing page are withdrawn and replaced with correction notices at articles/gemini-3-5-pro-review and pricing/gemini-3-5-pro; every other mention across the site has been corrected or removed.

GPT-5.6 availability corrected twice. The July 8 correction correctly walked back benchr's premature July 1 GA claim. OpenAI then confirmed the actual API release on July 9, 2026 in its official changelog, so the current status is GA. Pricing stayed $5/$30, $2.50/$15, and $1/$6; the current docs list 1.05M context and 128K max output for all three tiers.

Claude Sonnet 5 max output corrected 200,000 → 128,000 tokens. Re-read against Anthropic's Models overview docs: Sonnet 5's max output is 128,000 tokens, not 200,000 as recorded at launch. That figure ties Claude Opus 4.8 and is double Sonnet 4.6's 64,000-token limit. Corrected in model-figures.json, models.json, the Sonnet 5 launch and pricing pages, and every comparison that cited the old figure.

New deprecation: DeepSeek's legacy deepseek-chat and deepseek-reasoner aliases. Per api-docs.deepseek.com, these legacy aliases retire July 24, 2026 in favor of the explicit versioned names (DeepSeek V4-Flash / V4-Pro) — an alias rename only, no pricing or capability change. Added to deprecations.json, the deprecations hub, and the tracker.

Full writeups of the three corrections above are logged at /corrections.

Pricing correction
Claude Sonnet 5 pricing corrected after re-reading Anthropic's official pricing page

Corrected Claude Sonnet 5 pricing in model-figures.json, models.json, history.json, and the English/Arabic Sonnet 5 launch and pricing pages. Anthropic's official rate is $2/1M input and $10/1M output through August 31, 2026, then $3/$15 from September 1. The July 1 benchr entry previously recorded $4/$20 incorrectly.

New model
GPT-5.6 preview system card updated — Sol added to the ranked index; Terra and Luna's full specs recorded (still GA)

OpenAI's GPT-5.6 system card was updated on July 1, 2026 with new benchmark scores and confirmed specs while access was still limited to preview partners. The July 9 API changelog later confirmed GA. Sol is benchr's canonical models.json entry: $5/$30 per 1M, cached input $0.50, 1.05M context, 128K max output, API id gpt-5.6-sol. Terra and Luna are tracked in model-figures.json with official prices $2.50/$15 and $1/$6 and the same 1.05M context / 128K output limits.

New models
Claude Sonnet 5 added to the index; Gemini 3.5 Pro entry retracted as fabricated

Correction, July 8, 2026: this entry originally added a "Gemini 3.5 Pro" model that Google never released — the claim was fabricated, and the addition has been retracted. See the July 8, 2026 entry above and /corrections for the full explanation.

Claude Sonnet 5 (released July 1, 2026) — Anthropic's second Mythos-class-architecture model after Claude Fable 5 — added to models.json (id claude-sonnet-5): 1,000,000 context, 128,000 max output (matching Claude Opus 4.8 and double Sonnet 4.6's 64,000), tentative retirement floor not sooner than July 1, 2027. The benchmark figures in this entry are provider-reported. Pricing corrected July 3: $2/1M input and $10 output through August 31, then $3/$15 from September 1; the earlier $4/$20 entry was wrong. New launch writeup and pricing page.

Availability
Claude Fable 5 restored to all customers; export-control review closes

The export-control review that suspended Claude Fable 5 for all customers on June 12 closed July 1, 2026: the Commerce Department lifted the suspension and AWS restored Bedrock access in step, the same day Anthropic launched Claude Sonnet 5. Claude Mythos 5 remains Project-Glasswing-only as before — unaffected by this restoration. No pricing or spec changes: $10/$50 per 1M, 1M context, 128K max output are unchanged. Recorded as a dated availability note on the Fable 5 entry in model-figures.json; updated launch writeup.

New model
OpenAI previews the GPT-5.6 series — Sol, Terra, and Luna — under government-limited access

On June 26, 2026 OpenAI began a limited preview of three GPT-5.6 models: Sol (flagship), Terra (GPT-5.5-class at roughly half the price), and Luna (cheapest and fastest), plus new max and ultra reasoning modes. Access was restricted to about 20 US-government-approved trusted partners via the API and Codex, with general availability planned “in the coming weeks.”

Announced pricing per 1M tokens was Sol $5/$30, Terra $2.50/$15, and Luna $1/$6. We recorded those rates in model-figures.json with a caveat: they came from OpenAI's preview announcement and help center but, as of June 28, were not yet on the machine-readable developer pricing page. Context, maximum output, and exact API IDs were also unpublished at that point.

We added a GPT-5.6 launch record and pricing page. The models stayed out of the interactive ranking until official benchmark data was available, the same rule applied to Kimi K2.7-Code.

Availability
Export-control update: Claude Mythos 5 partially restored; Claude Fable 5 still suspended

The June 12 U.S. export-control suspension of Anthropic's Mythos-class models eased on June 26: the Commerce Department cleared Claude Mythos 5 for redeployment to a defined set of US critical-infrastructure organizations — a limited Project Glasswing re-release, not general availability. Claude Fable 5, the generally available model, remains fully suspended for all customers; the clearance did not cover it, and its $10/$50 listed pricing is unchanged.

The same June 2 “covered frontier model” Executive Order is the backdrop for OpenAI's government-gated GPT-5.6 release. Recorded on the Fable 5 and Mythos 5 entries in model-figures.json; background on the export-control explainer.

Deprecation
Gemini 3 Pro Image Preview and Gemini 3.1 Flash Image Preview are now retired

Both image-preview models reached their June 25, 2026 shutdown — re-confirmed on Google's official Gemini deprecations table. They were flipped deprecated → retired in deprecations.json (the hub moves them into the retired record); migrate to the GA gemini-3-pro-image and gemini-3.1-flash-image. Still upcoming: Claude Mythos Preview retires June 30, 2026.

Data
Five models added to the verified record; Qwen and Grok prices corrected

Added to model-figures.json: Mistral OCR 4 (per-page $4/$2/$5 per 1,000 pages), Mistral Small 4 ($0.15/$0.60, 119B/6B MoE, 256K), Qwen3.7-Max (hosted, $2.50/$7.50), Grok Build 0.1 (agentic coding, $1.00/$2.00, 256K), and the open-weight Qwen-AgentWorld 35B-A3B and 397B-A17B.

Corrections from the same pass: Qwen3.7-Plus and Qwen3.6-Plus now carry official per-token prices on Alibaba Cloud Model Studio (resolving the prior “not published” null) — tiered $0.4–$1.2 / $1.6–$4.8 and $0.5–$2 / $3–$6 respectively; and Grok 4.20 now lists $1.25/$2.50 at 1M context on docs.x.ai, reconciled from the $2/$6, 2M launch figure. No commercial price changed at any other tracked provider since June 23.

New model
Kimi K2.7-Code added to the verified record, with a pricing page

Moonshot's coding-specialist successor to K2.6 is now in model-figures.json: Kimi K2.7-Code at $0.95/1M input (cache miss), $0.19 cache hit, $4.00 output, 256K context, Modified MIT — plus a Highspeed tier at double the rates ($1.90/$0.38/$8.00) for ~180 tok/s. Both verified June 23 on platform.kimi.ai and the official Hugging Face model card; Moonshot publishes no official release date. New Kimi K2.7-Code pricing page, and the Kimi review now carries an update section. Held out of the ranking tools for now — no official benchmarks to score it on.

Verification
GPT-5.5 Pro pricing now official ($30/$180); Mistral date and hosted Qwen flagships recorded; no tracked price changed

A full re-check against provider pages found no commercial price change at any tracked provider since June 12. GPT-5.5 Pro standard pricing is now readable on developers.openai.com — $30/1M input, $180/1M output (batch $15/$90, no cached-input discount) — so the previously-null fields in model-figures.json are filled, and the GPT-5.5 pricing page notes the tier.

Corrections from the same pass: Mistral Medium 3.5's official release date is April 28, 2026; the hosted Qwen flagships Qwen3.6-Plus (1M context), Qwen3.6-Max-Preview, and Qwen3.7-Plus (multimodal) are confirmed on Alibaba Cloud — per-token pricing isn't officially published, so it stays null.

Deprecation
OpenAI deprecates older GPT-5 and o3 snapshots — API removal December 11, 2026

Logged OpenAI's June 11, 2026 notice: the dated snapshots gpt-5-2025-08-07, gpt-5-mini/nano-2025-08-07, gpt-5-pro-2025-10-06, o3-2025-04-16, and o3-pro-2025-06-10 are removed from the API December 11, 2026. The floating gpt-5 / gpt-5-mini / gpt-5-nano aliases remain available. Added to deprecations.json, the deprecations hub, and the tracker. Imminent and unchanged: Gemini 3 Pro Image Preview (June 25), Claude Mythos Preview (June 30).

Deprecation
Claude Sonnet 4 and Claude Opus 4 are now retired

Their June 15, 2026 shutdown has passed — re-confirmed on Anthropic's official model-deprecations table. Calls to claude-sonnet-4-20250514 and claude-opus-4-20250514 now fail. Both were flipped deprecated → retired in deprecations.json; the deprecations hub moved them into the retired record, and the Sonnet 4 and Opus 4 migration pages now read as in-effect. Opus 4.1 remains deprecated until August 5, 2026.

Availability
Claude Fable 5 and Mythos 5 suspended for all customers (U.S. export-control order)

On June 12 a U.S. Commerce Department export-control directive barred Mythos-class models from use by any foreign national, and Anthropic suspended both Claude Fable 5 and Claude Mythos 5 for all customers; AWS revoked Bedrock access the same day. All other Claude models are unaffected.

Listed pricing and specs are unchanged in model-figures.json (a dated availability note was added); Anthropic disputes the recall and says it is working to restore access. Coverage updated on Claude Mythos, the Fable 5 launch, and the Fable 5 pricing, comparison, and model-picker pages.

New section
AI API Error Database launched — 15 errors, 3 providers

New /errors reference: 15 common AI API errors across OpenAI, Anthropic, and Gemini, each verified against the provider's official error documentation, with causes, code fixes, and migration bridges into the deprecations record and pricing pages. Structured dataset at api-errors.json; validator at scripts/generate_error_pages.py.

Verification
Full price re-verification (no changes found); cache-hit prices and retirement floors added; 9 deprecation entries logged

Every tracked commercial price was checked against the provider's own pricing page — no price changes since June 10 (Anthropic, OpenAI, Google, xAI, DeepSeek, Moonshot, Mistral). Anthropic's pricing page now publishes per-model cache-hit prices, closing previously-honest gaps in model-figures.json: Opus 4.8/4.7 $0.50, Sonnet 4.6 $0.30, Haiku 4.5 $0.10, Fable 5/Mythos 5 $1 per 1M cached input.

Status correction: Claude Opus 4.6 is officially Active (the entry previously called it legacy), with fast mode at $30/$150. Anthropic's official "not sooner than" tentative retirement floors logged for all active Claude models (e.g. Opus 4.8 → May 28, 2027; Sonnet 4.6 → February 17, 2027).

Deprecation record expanded by 9 entries from a full re-read of provider docs: OpenAI's July 23 wave (five Codex IDs, o4-mini-deep-research, computer-use-preview), GPT-4.1 nano and GPT Image 1 joining the October 23 list (now eleven IDs), the November 30 platform retirements (Prompts API, Evals, Agent Builder — announced June 3), and Gemini 3 Pro Image Preview (June 25). Details on the deprecations hub. Gemini 3.5 Pro remains unreleased — nothing official to add.

New section
Deprecations hub + price-history record launched

New /deprecations hub: every announced retirement across Anthropic, OpenAI, and Google with shutdown dates, replacements, and migration price math — plus five dedicated migration pages (Claude Sonnet 4, Claude Opus 4/4.1, GPT-4o, OpenAI's October wave, Gemini 2.5), a deprecations RSS feed, and an open deprecations.json dataset. All entries verified June 12 against official provider deprecation docs.

New /price-history page: the append-only record of verified AI API pricing events, published as open data (CC BY 4.0) from history.json.

Tracker corrections from the same verification pass: GPT-4o and GPT-4 Turbo retire October 23, 2026 (the tracker previously showed early-2026 dates); Gemini 3 Pro Preview's API shutdown was March 9, 2026 (not March 26); Claude 3 Opus (Jan 5) and Claude 3.5 Haiku (Feb 19) are now marked retired; Claude Mythos Preview (June 30) and Gemini 2.5 Pro/Flash (October 16) added.

New models
Claude Fable 5 added; GPT-5.4 re-added after verification — index now 21 models

Claude Fable 5 (released June 9, 2026) added to the index: $10/$50 per 1M, 1M context, 128K output, API id claude-fable-5. First Mythos-class model in general availability. Source: anthropic.com/news/claude-fable-5-mythos-5.

GPT-5.4 re-added. It was removed on June 1 as "unverified" — that call was wrong. Re-verified June 10 against OpenAI's pricing page and release notes: released March 5, 2026, $2.50/$15 per 1M, cached input $0.25, context up to 1M. The same pass confirmed Claude Opus 4.6 (Feb 5), Grok 4.20 (Mar 10, $2/$6, 2M context), and Qwen 3.5 (Feb 16) are all real; entries added to model-figures.json.

Tracker updated: Claude Sonnet 4 and Claude Opus 4 retire June 15, 2026; Claude Opus 4.1 retires August 5, 2026 (per the official Anthropic docs).

New models
Tool suite launch — 5 phases, 19 models

Launched the full benchr tool suite: ranked index, charts, cost calculator, model recommender, deprecation tracker, and release timeline. All tools read from models.json — never duplicate data per page.

New models added to the index: Claude Opus 4.8, GPT-5, GPT-5 Mini, Grok 4.3, Kimi K2.6, Mistral Large 3, Llama 4 Scout, Llama 4 Maverick. Removed unverified entries (GPT-5.4, GPT-5.4 mini, Qwen 3.5). All 19 current models verified against official provider pricing pages.

New model
Claude Opus 4.8 added
Released May 28, 2026. $5/$25 per 1M standard; $10/$50 Fast Mode. SWE-Bench Pro 69.2%. Source: anthropic.com/news. Review →
New model
Gemini 3.5 Flash (GA) added
GA since May 19, 2026. $1.50/$9.00 per 1M; free API tier; 1,048,576-token context and 65,536-token max output. Google does not publish a universal throughput figure. Source: ai.google.dev. Review →
Pricing fix
GPT-5.5 context corrected to 1,048,576 tokens
Previous models.json had 1M context for GPT-5.5. Corrected to 1,048,576 per OpenAI's API documentation. Source: platform.openai.com/docs.
Data correction
Removed unverified model entries: GPT-5.4, GPT-5.4 mini, Qwen 3.5
These entries couldn't be confirmed against official provider sources. Replaced with verified models: GPT-5 ($1.25/$10, Aug 2025) and GPT-5 Mini ($0.25/$2.00). Qwen 3.5 entry removed; Qwen 3.6 remains (verified Apache 2.0, Apr 2026).
New
benchr Rating methodology documented
Introduced the benchr Rating: capability × 0.65 + price efficiency × 0.35. Formula runs in open JavaScript at assets/js/models.js. Full methodology documented on the rankings page. No provider pays for ranking position.
Data correction
GPT-5 pricing corrected from $2.50/$15 to $1.25/$10
Prior data in gpt-5-5-review.html incorrectly stated GPT-5 at $2.50/$15. Actual price confirmed at $1.25/$10 from official OpenAI pricing page. Source: openai.com/pricing.
Deprecation noted
Gemini 3 Pro sunset confirmed (March 26, 2026)
Google confirmed Gemini 3 Pro was sunset March 26, 2026. Successor: Gemini 3.1 Pro. Added to deprecation tracker. Source: ai.google.dev/deprecation.

Raw snapshot history (append-only): assets/data/history.json. Spot an error? File a correction · Contact· Privacy · Terms.

→ Deprecation tracker → Release timeline → Full model rankings → models.json (raw data)
Updates

Follow new pieces through RSS, recent releases, and the changelog.