Claude Sonnet 5 launches: Mythos-class architecture at Sonnet pricing

Anthropic's second Mythos-class model is not a new flagship. It launches at $2/$10 through August 31, then $3/$15 from September 1, with a 128K max output and a SWE-bench Verified score that edges out last month's Opus 4.8.

By benchr Editorial Team · · · View changelog · Availability, max output, and refusal behavior corrected against current records

Claude Sonnet 5 launches: Mythos-class architecture at Sonnet pricing: warm clay layers and branching decision routes.
Benchr model field plate Claude Sonnet 5 Two tiers · shared architecture
AnthropicClaude Sonnet 5 launches: Mythos-class architecture at Sonnet pricing is mapped with warm clay layers and branching decision routes.
Intro input / 1M $2 Output $10 through August 31; then $3/$15
Context window 1M Same as Sonnet 4.6 and Opus 4.8
Max output 128K Sonnet 4.6: 64K. Opus 4.8: 128K
SWE-bench Verified 89.4% Edges out Opus 4.8's 88.6%

Three weeks ago, Anthropic's Mythos-class architecture lived in exactly one place: Claude Fable 5, priced at $10 per million input tokens and $50 per million output for the hardest agentic work. The open question was whether that architecture would ever come down in price, or stay a flagship-only luxury. On July 1, Anthropic answered it. Claude Sonnet 5 runs on the same Mythos-class architecture, launches at $2 per million input tokens and $10 per million output through August 31, then settles at $3/$15 from September 1 — not a new "Sonnet 4.7," but a second model built on Fable 5's foundation at the Sonnet price tier.

One architecture, two tiers

Sonnet 5 launches below Sonnet 4.6 during the introductory window: $2 per million input tokens and $10 per million output through August 31, 2026. On September 1 it moves to the standard Sonnet rate of $3/$15, matching Sonnet 4.6 and staying below Opus 4.8's $5/$25. Cached input is $0.20 per million during the intro period and $0.30 afterward; Batch cuts both directions in half to $1/$5 during intro and $1.50/$7.50 after. Context stays at 1M tokens, matching Sonnet 4.6 and Opus 4.8, but max output jumps to 128,000 tokens, double Sonnet 4.6's 64,000 and matching Opus 4.8's 128,000. The API id is claude-sonnet-5, and Anthropic's tentative retirement floor is not sooner than July 1, 2027.

The architecture brings two more inherited traits. Adaptive thinking is always on, with no extended-thinking toggle to flip — the same behavior Fable 5 introduced. The same safety classifiers also apply: classified offensive-cyber, biology/chemistry, or distillation requests return an explicit refusal. Another-model retry occurs only when configured by the application or product.

What the benchmarks say

The headline number is SWE-bench Verified: Sonnet 5 scores 89.4%, ahead of Claude Opus 4.8's 88.6% — a mid-tier model beating last month's flagship on a closely watched metric. SWE-bench Pro, the harder agentic-coding test, comes in at 71.8%. Terminal-Bench 2.1 lands at 85.6%. Reasoning tells a different story: GPQA Diamond is 92.0%, behind Opus 4.8's 93.6%, so Opus 4.8 keeps the edge on graduate-level science reasoning. On ARC-AGI-2, Sonnet 5 scores 20.0, ahead of Sonnet 4.6's 15.0. Humanity's Last Exam without tools comes in at 42.5%. The rest of the sheet: LMSYS Arena 1435, MMLU 93.8%, HumanEval 96.0%, MATH 93.5%.

Read the pattern honestly: Sonnet 5 closes almost all of the coding gap to Opus 4.8, and actually passes it on SWE-bench Verified, while giving up ground on the hardest reasoning benchmark. At the corrected launch price, that is a sharper trade: half of Opus 4.8's input rate during the introductory window, then 60% of Opus 4.8's input rate from September 1.

Sonnet 5 vs Sonnet 4.6 vs Opus 4.8

The Claude line after July 1, 2026, from the official docs
SpecClaude Sonnet 5Claude Sonnet 4.6Claude Opus 4.8
Price (in/out per 1M)$2 / $10 intro; $3 / $15 from Sep. 1$3 / $15$5 / $25
Context window1M tokens1M tokens1M tokens
Max output128K64K128K
SWE-bench Verified89.4%79.6%88.6%
GPQA Diamond92.0%89.9%93.6%
Thinking modeAdaptive, always onStandardStandard
RestrictionsClassified cyber / bio / distillation requests refuse explicitlyStandardStandard

Is the mid-tier upgrade worth it?

Run the math on a real workload before switching. A coding agent burning 2M input tokens and 400K output tokens a day costs about $8 on Sonnet 5 during the intro window (2 × $2 + 0.4 × $10), against $12 on Sonnet 4.6 (2 × $3 + 0.4 × $15) and $20 on Opus 4.8 (2 × $5 + 0.4 × $25). From September 1, the same Sonnet 5 workload rises to $12, matching Sonnet 4.6 while keeping the 128K max-output ceiling and the stronger coding score. The cost calculator will run this against your own volumes, and the Claude Sonnet 5 pricing breakdown covers the caching and batch math in full.

Where the upgrade clearly pays off: coding agents and long-running tool loops that were hitting Sonnet 4.6's 64K output ceiling, since 128K max output means far fewer truncated responses mid-task. Where it is a harder sell: workloads that are already comfortable on Sonnet 4.6 and do not need the extra output headroom or the coding bump — once the intro window ends, the price match is nice, but it is still a migration to test rather than a blind swap. And if your work depends on Opus 4.8's GPQA-level reasoning ceiling, Sonnet 5 does not close that gap; Opus 4.8 stays the pick.

A crowded launch week

Sonnet 5 didn't ship in isolation. The same day, Anthropic's export-control review closed and Claude Fable 5 was restored to all customers, with AWS reinstating Bedrock access. OpenAI's GPT-5.6 was partner-gated at that point and later reached general availability on July 9. None of that changes the Sonnet 5 math directly, but it establishes the current availability timeline.

Frequently asked

What is Claude Sonnet 5?

Claude Sonnet 5 is Anthropic's second Mythos-class-architecture model, launched July 1, 2026, priced at $2/$10 per million tokens through August 31 and $3/$15 from September 1. It has a 1M context window and 128,000-token max output. Adaptive thinking is always on; safety-classified requests return an explicit refusal, and another-model retry requires configured application logic.

How much does Claude Sonnet 5 cost?

$2 per million input tokens and $10 per million output through August 31, 2026. Starting September 1, standard pricing is $3/$15. Cached input is $0.20 per million during intro and $0.30 afterward; Batch halves both directions.

How does Sonnet 5 compare to Sonnet 4.6 and Opus 4.8?

SWE-bench Verified: 89.4% for Sonnet 5, ahead of Sonnet 4.6's 79.6% and narrowly ahead of Opus 4.8's 88.6%. Opus 4.8 still leads GPQA Diamond, 93.6% to 92.0%. Sonnet 5's 128K max output beats Sonnet 4.6's 64K and matches Opus 4.8's 128K.

Is Sonnet 5 worth upgrading to from Sonnet 4.6?

Test it if you need the extra output headroom or want to compare its provider-reported coding result with Sonnet 4.6 and Opus 4.8. Keep Opus 4.8 in the set for deep reasoning, and design explicitly for safety-classified requests that refuse rather than routing automatically.

Changelog

  • July 23, 2026 — Corrected GPT-5.6's current status to July 9 general availability and replaced the unsupported automatic-Opus fallback claim with explicit refusal plus configured retry behavior.
  • July 1, 2026 — Published. Pricing, context window, max output, benchmark figures, classifier behavior, and the retirement floor verified against Anthropic's launch announcement and the official model docs.

References

  1. Anthropic, "Claude Sonnet 5," anthropic.com/news/claude-sonnet-5, July 1, 2026. Source for the release, classifier behavior, and benchmark figures; pricing corrected July 3 against Anthropic API pricing docs.
  2. Anthropic, "Pricing," platform.claude.com/docs/en/about-claude/pricing, re-verified July 3, 2026. Source for the corrected price.
  3. Anthropic, "Models overview," platform.claude.com/docs, accessed July 1, 2026. Source for the API id, context window, max output, and the retirement-floor policy.
  4. Anthropic, "Claude Opus 4.8 system card," anthropic.com. Source for the Opus 4.8 comparison figures.