Archive
Every piece, ordered by the period it covers. benchr launched on May 30, 2026 — items listed under earlier months are retrospective coverage, and each article's dateline shows its actual publication date. The summaries identify each article's scope: provider figures are attributed inside the article, while model-fit language is an editorial hypothesis to reproduce on your own workload, not an unpublished benchr result. Index reviewed August 21, 2026. Browse AI selection guides by use case, or subscribe via RSS.
August 2026
21 Aug·26
Gemini 3.7 Flash: price the introductory window before you migrate.
Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.
21 Aug·26
Grok 4.6: a 500K agent model whose output ceiling is not a number.
xAI targets long coding runs and interactive agents at $2/$6, while its docs publish no numeric text-output cap.
21 Aug·26
GLM-5.3: a 1M coding model with weights still behind a safety gate.
Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.
21 Aug·26
Kimi K3: open weights at 2.8T parameters, but only 104B active.
Moonshot's sparse agent model pairs a 1M context with flat hosted pricing and a deployment decision bigger than the token rate.
-
14 Aug·26
DeepSeek's old API aliases are retired: a migration checklist.
The legacy chat and reasoner aliases stopped being usable on July 24. Replace them with an explicit V4 ID and test the full request contract.
-
08 Aug·26
Grok 4.5, reviewed: the coding specialist that got smaller, not bigger.
Built alongside Cursor for coding and agent work, priced above Grok 4.3, and running a context window that shrank rather than grew.
-
08 Aug·26
GLM-5.2, reviewed: an open-weight model with a real coding claim.
Zhipu's MIT-licensed 753B model undercuts GPT-5.5 on price by more than 20x and claims to beat it on several coding benchmarks.
-
04 Aug·26
OpenAI Presence is an enterprise agent service, not another API.
What the managed voice-and-chat product includes, how to read OpenAI's own outcome figures, and the commercial questions buyers still need answered.
July 2026
28 Jul·26
Claude Opus 5: the migration question behind Anthropic's new flagship
Opus 5 keeps the $5/$25 price and 1M-token window. The catch is a thinking setting that can break an otherwise routine upgrade.
28 Jul·26
GPT-Realtime-2.1: evaluate the conversation loop, not a text-only price card
This voice model is built around interruptions, tools, and recovery. Its audio tokens also make the text price a poor budget shortcut.
28 Jul·26
GPT-Realtime-2.1 mini: the lower-cost voice model still has an audio bill
Mini keeps the 128K window and cuts every listed rate. Audio still dominates the calls you need to measure.
28 Jul·26
GPT-Live-1 is a ChatGPT Voice rollout, not an API endpoint
GPT-Live-1 changes ChatGPT Voice. It does not yet give developers an endpoint, a model ID, or an API price to build around.
28 Jul·26
GPT-Live-1 mini: the Voice fallback is not a developer SKU
The smaller Live model is part of ChatGPT's Free-tier experience. It is not a public API SKU, and “mini” is not a price.
28 Jul·26
Gemini 3.5 Flash-Lite: the $0.30 subagent target has a migration cost
Flash-Lite's $0.30 input rate suits large queues. Before switching, check the thinking defaults and request fields Google changed.
28 Jul·26
Gemini 3.1 Flash TTS Preview: steerable speech with preview boundaries
Google publishes the input and output limits and supports Batch. The model card still leaves the price blank.
28 Jul·26
Gemma 4 26B A4B IT: an official availability record, not a filled-in spec sheet
Google published the exact model ID and where to find it. Context, pricing, and benchmark figures are still missing.
28 Jul·26
Gemma 4 31B IT: a public model name without a public deployment spec
The 31B IT model is public, but its deployment details are not. Do not let the name stand in for limits, pricing, or results.
28 Jul·26
Leanstral 1.5 is for proofs, not chat: read its benchmark claims in context
Leanstral is built for formal proofs. Its open weights and proof scores do not make it a general chat model.
28 Jul·26
Qwen-Audio 3.0 TTS Plus: built-in voices are the product boundary
Plus is the built-in-voice branch of Qwen's new TTS line. Alibaba has not published token limits or a per-token price.
28 Jul·26
Qwen-Audio 3.0 TTS Flash: latency and cloning change the deployment decision
Flash adds low-latency synthesis and voice cloning. That makes consent and end-to-end call testing part of the model choice.
28 Jul·26
Muse Image: Meta's consumer launch leaves a developer checklist blank
Muse Image is live inside Meta AI. Developers still have no public model ID, rate card, or integration terms.
28 Jul·26
Muse Video Preview: native audio is a preview feature, not an API contract
Muse Video can generate native audio, but Meta has only previewed it. There is no public API contract to compare yet.
28 Jul·26
Correction: Gemini 3.5 Pro has not been released
The earlier review was wrong. This permanent correction explains what was retracted and points readers to Google's actual current records.
-
22 Jul·26
Gemini 3.6 Flash launch: cheaper output, same Flash input.
Google's new stable Flash model keeps the $1.50 input price, cuts output to $7.50, and becomes the cleaner migration target for older Flash endpoints.
-
1 Jul·26
Claude Sonnet 5 launches: Mythos-class architecture at a mid-tier price.
$2/$10 introductory pricing, a 128K maximum output, and an Anthropic-reported SWE-bench Verified score above its published Opus 4.8 figure; neither score is a benchr reproduction.
June 2026
-
28 Jun·26
GPT-5.6: OpenAI ships Sol, Terra, and Luna behind a government gate.
A new frontier series in limited preview to about 20 approved partners. Announced at $5/$30, $2.50/$15, and $1/$6 — with context windows and API IDs still unpublished.
-
19 Jun·26
When a large open model does not fit your computer.
A hardware-sizing guide to quantization, offloading, smaller checkpoints, and cloud GPUs, with measurements to reproduce on your planned configuration.
-
19 Jun·26
How to fine-tune an open model on your own data without owning a GPU.
A reproducible QLoRA workflow for rented GPUs, with dated cost-planning scenarios rather than a universal training-price promise.
-
19 Jun·26
“CUDA out of memory”: why it happens, and how to run a model that's too big for your card.
A diagnostic sequence for common GPU-memory failures, from configuration changes to a larger card, with the required memory checks made explicit.
-
19 Jun·26
Renting a GPU vs. paying per token: a break-even worksheet.
Compare GPU dollars per hour with API dollars per million tokens using your measured utilization, throughput, operations time, and review burden.
-
16 Jun·26
The U.S. pulled Claude Fable 5 and Mythos 5.
The dated June suspension, the July 1 restoration, and what the interruption teaches about model continuity.
-
13 Jun·26
Which Claude model should you use in 2026?
Four models you can buy, one you can't. A plain decision guide to Anthropic's lineup, by task and by budget.
-
13 Jun·26
Claude Fable 5 vs GPT-5.5 vs Gemini 3.1 Pro.
The three newest frontier models, head to head on price, context, the benchmarks each lab published, and which one fits which job.
-
13 Jun·26
Claude Opus 4.8 vs Gemini 3.1 Pro.
Two frontier candidates compared using each provider's published benchmarks and prices. Treat workload fit as a hypothesis and reproduce it under one local rubric.
-
10 Jun·26
GPT-5.4, reviewed as a value candidate.
A dated price-and-capability review that turns the value claim into a workload-specific cost and quality check.
-
10 Jun·26
Claude Fable 5 is the Mythos-class model you can finally use.
$10/$50, 1M context, explicit safety refusals, and the current post-restoration access terms.
-
09 Jun·26
Claude Opus 4.8 is live, and the pressure is on coding and agents.
Anthropic's new flagship lands at $5/$25, with a fast mode and a Mythos-class security model arriving in the coming weeks.
-
09 Jun·26
Anthropic's Project Glasswing and the case for Claude as a security tool.
An expansion to ~200 critical-infrastructure orgs, and what Claude Security and the Mythos preview do.
-
09 Jun·26
GPT-5.5 pricing, and the 272K-token cliff that doubles your bill.
Standard rates, the long-context reprice that hits the whole session, and a cost worksheet for deciding whether GPT-5.5 Pro clears your threshold.
-
09 Jun·26
Google's developer AI at I/O 2026: from autocomplete to agents.
Gemini 3.5 Flash as the agent default, the Gemini CLI to Antigravity migration on June 18, and what breaks.
-
09 Jun·26
Grok 4.3 is now xAI's default, and old slugs bill at its prices.
Legacy model aliases now redirect to Grok 4.3 and are billed at its rates. What that means for your API costs.
-
09 Jun·26
Google's AI Search is becoming an agent layer, not a summary box.
Background information agents and agentic booking, set against the click-loss data from Pew, Ahrefs, and Seer.
-
09 Jun·26
The UK just forced Google to give publishers AI Search control.
The UK CMA's published opt-out requirement and its stated relationship to ordinary search visibility, with scope and jurisdiction made explicit.
-
09 Jun·26
WebMCP and the push to make websites agent-readable.
A proposed standard from Google and Microsoft that lets your site hand agents a list of tools instead of making them guess.
-
06 Jun·26
DeepSeek vs OpenAI Pricing: Cost Comparison & Quality Trade-offs.
A dated price comparison from official pages, plus provider-published coding benchmarks labeled by source; verify both on the live pages and your own tasks.
-
06 Jun·26
Anthropic Claude API Pricing Guide: Opus 4.8, Sonnet, & Haiku.
Complete Claude API pricing breakdown. Opus 4.8 at $5, Sonnet 4.6 at $3, Haiku 4.5 at $1 per million input tokens. Plus the 90% caching discount explained.
-
06 Jun·26
OpenAI API Pricing Guide: GPT-5.5, GPT-5, and GPT-5 Mini Costs.
GPT-5.5 at $5, GPT-5 at $1.25, GPT-5 Mini at $0.25 per million input tokens. Batch and caching discounts explained.
-
06 Jun·26
Low-cost LLM APIs in 2026: a dated price comparison.
Provider-listed sub-$1 input prices and self-hosted candidates, with output, caching, retries, and quality kept in the decision rather than a universal cheapest-model claim.
-
06 Jun·26
AI Model Pricing Comparison 2026: Cost per Million Tokens.
Complete comparison of API token pricing across OpenAI, Anthropic, Google, DeepSeek, and open-weights. Sourced from official docs.
May 2026
-
30 May·26
Claude Mythos 5: invite-only through Project Glasswing.
Access returned July 1, but remains restricted to approved Glasswing organizations without self-service signup.
-
30 May·26
Claude Cowork: the desktop agent that isn't for coders.
Give it a goal, point it at your files, let it work. Claude Code's engine, aimed at everyone who isn't a coder.
-
30 May·26
GPT-5.5, reviewed: is the upgrade off GPT-5 worth it.
OpenAI put the gains into agentic coding and computer use, at roughly double GPT-5's API price. Who moves, who waits.
-
30 May·26
Gemini 3.1 Pro, reviewed.
Google-published reasoning results, the long-context price, and a checklist for reproducing the claimed improvement on your workload.
-
30 May·26
Gemini 3.5 Flash, reviewed.
A provider-priced candidate for agent loops, with output speed and end-to-end cost left as measurements to reproduce rather than asserted results.
-
30 May·26
Grok 4.3, reviewed.
xAI's documented live-web and X access, plus a test plan for deciding when fresh context helps enough to offset citation and reliability risks.
-
30 May·26
DeepSeek-V4, reviewed.
An MIT-licensed downloadable model with DeepSeek-published coding results; reproduce quality and total serving cost before comparing it with paid APIs.
-
30 May·26
Qwen3.6, reviewed.
Qwen documents a free Apache-licensed family in two sizes; shortlist by memory and workload, then reproduce quality before deployment.
-
30 May·26
Kimi K2.6, reviewed.
An open-weight trillion-parameter model that runs a swarm of sub-agents across thousands of steps. Free to download, cheap on the API.
-
30 May·26
Llama 4, reviewed.
A 10-million-token context on open weights still turns heads. But Meta has moved on, and Llama 4 is the last open Llama.
-
30 May·26
Mistral Large 3, reviewed.
Mistral's published parameter scale and Apache-2.0 license, plus the hardware and quality checks required before deployment.
-
30 May·26
ChatGPT Images 2.0, reviewed.
A reproducible image-text and dense-layout evaluation plan for GPT Image 2, without claiming an unpublished first-place result.
-
30 May·26
Opus 4.8 vs GPT-5.5: the coder's flagship vs the daily driver.
Both list $5 per million input tokens; compare output price and provider-published capability claims, then reproduce workload fit under the same rubric.
-
30 May·26
Gemini 3.1 Pro vs GPT-5.5: reasoning vs knowledge work.
These two flagships aim at different scoreboards. One chases hardest-mode reasoning, the other all-round professional work. Picking between them starts with that.
-
30 May·26
Grok 4.3 vs ChatGPT: when live context matters.
Compare documented live-web access, source traceability, general task coverage, latency, and cost on a dated question set.
-
30 May·26
ChatGPT vs Claude vs Gemini: the everyday pick for 2026.
Three subscriptions with different default models, compared through a workload worksheet rather than a universal subscription winner.
-
30 May·26
Claude vs ChatGPT for long-form writing.
Before voice or style, one boring number decides a lot: how much can each model write in one pass?
-
30 May·26
AI search engines compared: Perplexity vs ChatGPT Search vs Google AI.
An AI answer is only as good as your ability to check it. The question is which one shows its work.
-
30 May·26
Free coding-model candidates: DeepSeek vs Qwen vs Kimi.
Three downloadable or free-access families shortlisted from provider-published scores, followed by a held-out repository test plan.
-
30 May·26
AI video tools in 2026: an evaluation guide.
Compare current availability, prompt fidelity, motion, audio, failure rate, rights, and cost without assuming a universal leader.
-
30 May·26
AI tools for social media: a selection guide.
A platform-by-platform checklist for captions, hooks, repurposing, scheduling, approvals, and current plan limits.
-
30 May·26
Free AI coding tools: a plan comparison.
Documented $0 limits from Copilot, Cursor, and peers, plus a repository test plan and the upgrade triggers to track.
-
30 May·26
AI for long-form writing: an evaluation guide.
Blind-score voice, factual fidelity, revision effort, and continuity on your own held-out drafts; no private ranking is asserted.
-
30 May·26
AI study tools: a responsible-use guide.
Studying, summarizing, and problem-solving, with policy, privacy, citation, learning-value, and free-tier checks.
-
30 May·26
AI for resumes and cover letters: a selection guide.
Tailoring, ATS compatibility, factual accuracy, privacy, and application risks to validate before using a candidate.
-
30 May·26
AI for email, built-in or standalone.
Compare drafting, replies, retrieval, privacy, workflow fit, and price across built-in and standalone candidates.
-
30 May·26
AI for spreadsheets: a formula evaluation guide.
Test Excel Copilot, Gemini in Sheets, and chat models on a held-out workbook with formula, pivot, permission, and error checks.
-
30 May·26
AI research tools: a citation-checking guide.
Compare literature-search and summarization candidates by source coverage, traceability, quotation fidelity, and fabricated-reference rate.
-
30 May·26
AI for Arabic-English translation: an evaluation guide.
A reproducible bidirectional rubric for meaning, terminology, register, dialect, omissions, and long-document consistency.
-
30 May·26
AI for Saudi and Gulf Arabic: an evaluation guide.
A reproducible dialect rubric that measures Khaleeji retention, MSA fallback, Egyptian drift, terminology, and reviewer preference.
-
30 May·26
AI for customer service: a deployment guide.
Off-the-shelf resolution bots, platform agents, or build-your-own: what each costs and which fits your support volume.
-
30 May·26
Free AI tools with no subscription: a dated plan guide.
Provider-documented no-card options in 2026, with current limits, data terms, and upgrade triggers to verify before use.
-
30 May·26
The AI agent that checks out for you.
How agentic shopping works, who is building it, and where it can go wrong.
-
30 May·26
Are AI hallucinations fixed yet?
What got better by 2026, what did not, and the setups that cut made-up answers.
-
30 May·26
Which AI providers train on your chats.
Who learns from your conversations by default, how to opt out, and what stays private.
-
30 May·26
Do AI text detectors work?
The false-positive problem, who gets wrongly flagged, and what to do instead.
-
30 May·26
Do you need a reasoning model?
When the extra cost and latency of a thinking model pays off, and when it's wasted.
-
30 May·26
How to get cited inside AI answers.
GEO and AEO practices that may improve eligibility and source clarity, plus a measurement plan; no tactic guarantees a citation.
-
30 May·26
When the model remembers you.
How persistent memory works across chats, what it buys you, and the privacy trade.
-
30 May·26
What zero-click search did to the web.
Published studies on zero-click behavior, with their dates, samples, and limits rather than an unnamed private traffic result.
-
30 May·26
Claude Opus 4.8, reviewed.
Anthropic-published benchmark changes at the same listed price as 4.7, plus a bug-detection hypothesis to reproduce on held-out code.
-
30 May·26
Claude Sonnet 4.6, reviewed.
The $3/$15 daily-driver tier. When it's the right default, when to drop to Haiku, and when to pay for Opus.
-
30 May·26
Claude Haiku 4.5, reviewed.
The provider-listed $1/$5 Claude tier, with a held-out quality and retry-cost test for deciding when it meets your threshold.
-
30 May·26
Cutting your token bill.
Where AI token spend comes from, and the five levers that bring it down: routing, caching, batching, shorter output, lower effort.
-
21 May·26
Where popular AI benchmarks stop being decisive.
The documented limits of saturated or contaminated benchmarks, and a checklist for reading provider scores alongside local evaluation.
-
16 May·26
The limits of million-token context claims.
An editorial hypothesis comparing long windows with retrieval, plus a reproducible quality-and-cost test instead of a universal verdict.
-
11 May·26
Voice models compared: ElevenLabs, Whisper, OpenAI, Cartesia.
A provider-sourced voice-model guide with a reproducible worksheet for latency, Arabic narration, naturalness, consent, and deployment cost.
-
8 May·26
The price-per-use-case table.
What you pay for AI in 2026 by workload — chat, RAG, agents, batch — with five commercial models compared.
-
5 May·26
Prompt engineering did not die. It got narrower.
Three prompt techniques presented as hypotheses, with controlled before-and-after variants you can reproduce on held-out tasks.
-
4 May·26
AI for Arabic content: a working report on five models.
How Modern Standard, Saudi, Egyptian, and Levantine Arabic come out the other side of Claude, GPT-5, Gemini 3, Qwen 3, and Llama 4.
-
2 May·26
Multimodal evaluation plan: twelve images, four models.
A reproducible rubric for Claude, GPT-5, Gemini 3, and Llama 4; model placements are editorial hypotheses, not a private winner table.
April 2026
-
28 Apr·26
GPT-5 vs Claude Opus 4.7: a seven-task evaluation plan.
A refactor, a landing page, an obscure legal question, a recipe, a paper summary, a difficult email, and a broken script.
-
22 Apr·26
Claude Opus 4.7, reviewed.
Reproducible plans for repository refactoring, long-document review, multilingual checks, and a dated total-cost estimate.
-
17 Apr·26
RAG vs fine-tuning, with the math.
A reproducible quality and total-cost worksheet for choosing RAG, fine-tuning, long context, or a hybrid.
March 2026
-
18 Mar·26
AI agents, eighteen months in.
A bounded-deployment framework for LangGraph, OpenAI Assistants v2, Anthropic computer use, and AutoGen, with reliability hypotheses to test.
-
7 Mar·26
Running models on your own machine.
A reproducible local-throughput and total-cost plan across runtimes and quantizations, without unpublished device-speed claims.
-
1 Mar·26
Gemini 3 Pro, reviewed
Google-documented specifications plus workload hypotheses to validate on vision, long context, citations, latency, and cost.
February 2026
-
25 Feb·26
Small language models, in working use.
Phi-4, Gemma 3, and a held-out evaluation plan for deciding when a sub-10B model meets a bounded workload.
-
11 Feb·26
Context windows compared, across four frontier models.
A reproducible comparison of long context and retrieval across quality, omissions, citations, latency, and total token cost.
-
1 Feb·26
Coding assistants: a reproducible comparison plan.
A shared repository task, pinned versions, retained patches, automated tests, and blind review for Cursor, Copilot, Windsurf, and Cody.
January 2026
-
18 Jan·26
The open-weight tier right now: Llama 4, Mistral, Qwen, DeepSeek.
Provider-published open-model claims, license and deployment trade-offs, and a local rubric for comparing them with closed candidates.
-
4 Jan·26
GPT-5, reviewed.
OpenAI-documented specifications plus a reproducible plan for comparing speed, task breadth, and niche technical error rate.