Three weeks ago, Anthropic's Mythos-class architecture lived in exactly one place: Claude Fable 5, priced at $10 per million input tokens and $50 per million output for the hardest agentic work. The open question was whether that architecture would ever come down in price, or stay a flagship-only luxury. On July 1, Anthropic answered it. Claude Sonnet 5 runs on the same Mythos-class architecture, launches at $2 per million input tokens and $10 per million output through August 31, then settles at $3/$15 from September 1 — not a new "Sonnet 4.7," but a second model built on Fable 5's foundation at the Sonnet price tier.
One architecture, two tiers
Sonnet 5 launches below Sonnet 4.6 during the introductory window: $2 per million input tokens and $10 per million output through August 31, 2026. On September 1 it moves to the standard Sonnet rate of $3/$15, matching Sonnet 4.6 and staying below Opus 4.8's $5/$25. Cached input is $0.20 per million during the intro period and $0.30 afterward; Batch cuts both directions in half to $1/$5 during intro and $1.50/$7.50 after. Context stays at 1M tokens, matching Sonnet 4.6 and Opus 4.8, but max output jumps to 128,000 tokens, double Sonnet 4.6's 64,000 and matching Opus 4.8's 128,000. The API id is claude-sonnet-5, and Anthropic's tentative retirement floor is not sooner than July 1, 2027.
The architecture brings two more inherited traits. Adaptive thinking is always on, with no extended-thinking toggle to flip — the same behavior Fable 5 introduced. The same safety classifiers also apply: classified offensive-cyber, biology/chemistry, or distillation requests return an explicit refusal. Another-model retry occurs only when configured by the application or product.
What the benchmarks say
The headline number is SWE-bench Verified: Sonnet 5 scores 89.4%, ahead of Claude Opus 4.8's 88.6% — a mid-tier model beating last month's flagship on a closely watched metric. SWE-bench Pro, the harder agentic-coding test, comes in at 71.8%. Terminal-Bench 2.1 lands at 85.6%. Reasoning tells a different story: GPQA Diamond is 92.0%, behind Opus 4.8's 93.6%, so Opus 4.8 keeps the edge on graduate-level science reasoning. On ARC-AGI-2, Sonnet 5 scores 20.0, ahead of Sonnet 4.6's 15.0. Humanity's Last Exam without tools comes in at 42.5%. The rest of the sheet: LMSYS Arena 1435, MMLU 93.8%, HumanEval 96.0%, MATH 93.5%.
Read the pattern honestly: Sonnet 5 closes almost all of the coding gap to Opus 4.8, and actually passes it on SWE-bench Verified, while giving up ground on the hardest reasoning benchmark. At the corrected launch price, that is a sharper trade: half of Opus 4.8's input rate during the introductory window, then 60% of Opus 4.8's input rate from September 1.
Sonnet 5 vs Sonnet 4.6 vs Opus 4.8
| Spec | Claude Sonnet 5 | Claude Sonnet 4.6 | Claude Opus 4.8 |
|---|---|---|---|
| Price (in/out per 1M) | $2 / $10 intro; $3 / $15 from Sep. 1 | $3 / $15 | $5 / $25 |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | 128K | 64K | 128K |
| SWE-bench Verified | 89.4% | 79.6% | 88.6% |
| GPQA Diamond | 92.0% | 89.9% | 93.6% |
| Thinking mode | Adaptive, always on | Standard | Standard |
| Restrictions | Classified cyber / bio / distillation requests refuse explicitly | Standard | Standard |
Is the mid-tier upgrade worth it?
Run the math on a real workload before switching. A coding agent burning 2M input tokens and 400K output tokens a day costs about $8 on Sonnet 5 during the intro window (2 × $2 + 0.4 × $10), against $12 on Sonnet 4.6 (2 × $3 + 0.4 × $15) and $20 on Opus 4.8 (2 × $5 + 0.4 × $25). From September 1, the same Sonnet 5 workload rises to $12, matching Sonnet 4.6 while keeping the 128K max-output ceiling and the stronger coding score. The cost calculator will run this against your own volumes, and the Claude Sonnet 5 pricing breakdown covers the caching and batch math in full.
Where the upgrade clearly pays off: coding agents and long-running tool loops that were hitting Sonnet 4.6's 64K output ceiling, since 128K max output means far fewer truncated responses mid-task. Where it is a harder sell: workloads that are already comfortable on Sonnet 4.6 and do not need the extra output headroom or the coding bump — once the intro window ends, the price match is nice, but it is still a migration to test rather than a blind swap. And if your work depends on Opus 4.8's GPQA-level reasoning ceiling, Sonnet 5 does not close that gap; Opus 4.8 stays the pick.
A crowded launch week
Sonnet 5 didn't ship in isolation. The same day, Anthropic's export-control review closed and Claude Fable 5 was restored to all customers, with AWS reinstating Bedrock access. OpenAI's GPT-5.6 was partner-gated at that point and later reached general availability on July 9. None of that changes the Sonnet 5 math directly, but it establishes the current availability timeline.