Model intelligence · sourced Source → shortlist → decision
Open workspace →

Choose an AI model with evidence

Verified records, your tests, and real cost — one clear decision.

One clear model decision

Compare → test → choose.

Three steps. No universal winner.

33verified models
25logged changes
41retirement records
4 labelsevidence categories

New and changed

Recent model updates

View all changes →
LoadingReading the verified change record…
  1. 01

    DeepSeek raised prices and made the clock part of the bill

    DeepSeek raised V4 API prices on August 16, 2026 and split every rate into peak and off-peak halves. V4-Flash output went from $0.28 to $1.32 during peak hours.

  2. 02

    Gemini 3.7 Flash costs the same as 3.6 Flash — and both double in January

    Gemini 3.7 Flash went GA on August 13, 2026 at $0.75/$3.75 per 1M tokens, the same rate 3.6 Flash was cut to. Both return to $1.50/$7.50 on January 1, 2027.

  3. 03

    GLM-5.3: 1M coding, hosted now with weights pending

    Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.

  4. 04

    Gemini 3.7 Flash: introductory pricing and migration

    Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.

  5. 05

    Grok 4.6: 500K agent model with no numeric output ceiling

    Grok 4.6 targets long coding agents, starts at $2/$6, doubles at 200K prompt tokens, and has no numeric text-output cap.

  6. 06

    Grok 4.5, reviewed: the coding specialist that got smaller, not bigger

    Grok 4.5 review: xAI's coding specialist, built with Cursor, priced at $2/$6, with a smaller 500K context window than Grok 4.3 despite what some blogs claim.

  7. 07

    GLM-5.2, reviewed: an open-weight model with a real coding claim

    GLM-5.2 review: Zhipu's MIT-licensed, open-weight 753B model at $1.40/$4.40 per million tokens, with coding benchmarks Z.AI says beat GPT-5.5 and Opus 4.7.

  8. 08

    OpenAI Presence is an enterprise agent service, not another API

    OpenAI Presence bundles voice and chat agents with policies, evaluations, approved actions, and escalation. What buyers know—and what remains undisclosed.

Reviews

  1. 07

    Claude Opus 5

    Anthropic's current Opus at the same $5/$25 rate — what changes over 4.8, and who should move.

  2. 08

    GPT-5

    A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.

  3. 09

    Gemini 3 Pro

    A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.

Comparisons

  1. 10

    Coding assistants

    Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.

  2. 11

    Voice models

    ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.

  3. 12

    Context windows

    Advertised context capacity versus reproducible retrieval checks you can run on your own documents.

Analysis

  1. 13

    Price per use case

    The cheapest model for chat, coding, RAG, agents, classification, and summarization.

  2. 14

    The open-weight tier

    Where Llama 4, Mistral Large 3, DeepSeek-V4, and Qwen 3.6 stand against the closed labs.

  3. 15

    Why benchmarks stopped telling you

    MMLU is saturated above 90%. The benchmarks worth tracking now.

Guides

  1. 16

    Frontier models

    Claude Opus 4.7, GPT-5, and Gemini 3.1 Pro Preview: sourced specifications, editorial tradeoffs, and a local evaluation plan.

  2. 17

    Open-weight models

    Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.

  3. 18

    AI costs

    What AI costs by model and workload, and where teams overspend.

Compare models directly

An interactive comparison covering pricing, benchmarks, context windows, and capability ratings for all 33 models in the shared index, across frontier, mid, and open-weight tiers.

Open the comparison tool

All 33 models ranked by benchr Rating →

All 113 articles in the archive →

Recent model releases →

Provider guides: OpenAI, Anthropic, Google, and open weights →

API model ID directory and developer data →

benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.

Updates

Follow new pieces through RSS, recent releases, and the changelog.