Claude Sonnet 4.6 API pricing and evaluation context

Sonnet 4.6 is listed at $3/1M input and $15/1M output, with a 1M context window and 64K max output. Anthropic reports 79.6% on SWE-bench; that score does not establish production quality without a matched workload test.

By benchr Editorial Team · · Specifications and benchmark provenance rechecked against official sources, July 24, 2026 · View changelog

Input / 1MAnthropic
Output / 1MAnthropic
SWE-benchAnthropic-reported
Contextmax window

Head-to-head: how this rate card stacks against OpenAI's daily driver is covered in Sonnet 4.6 vs GPT-5, and against the open-weight challenger in DeepSeek V4-Pro vs Sonnet 4.6.

Pricing breakdown

claude-sonnet-4-6 — official Anthropic pricing
TierRate / 1M tokens
Standard input$3.00
Standard output$15.00
Cached input$0.30
Context window1,000,000 tokens
Max output64,000 tokens

How to test Sonnet against Opus

Anthropic's published scorecard reports 79.6% SWE-bench for Sonnet 4.6 and 88.6% for Opus 4.8. Those provider-reported scores describe a defined benchmark setup; the nine-point difference does not predict the acceptance-rate gap for chat, summarization, code review, or your repository.

Sonnet is a sensible lower-cost candidate. Run the same versioned prompts and tasks on both tiers, blind-review the outputs, and escalate only the task classes where Opus produces a measured improvement that justifies its higher token cost.

The 64K max output ceiling

Claude Sonnet 4.6 supports 64,000 output tokens per response. Claude Opus 4.8 supports up to 128,000 output tokens; both models list a 1M-token context window. A full 64K output at Sonnet pricing ($15/1M) costs $0.96, while the same token count at Opus 4.8 pricing ($25/1M) costs $1.60. Opus does not require splitting at 64K.

Caching and effective rates

Claude Sonnet 4.6 cached input costs $0.30/1M — 90% off the $3 standard rate. For a typical agent with a 50K-token system prompt running 1,000 calls per day, caching saves approximately $135/day versus uncached at $3 per million. With high cache hit rates, the effective input cost drops to roughly $0.57 per million — competitive with Haiku uncached pricing for input-heavy workloads.

Cost scenarios

At 20M input + 5M output per month: $60 + $75 = $135/month. Opus 4.8 at the same volume: $100 + $125 = $225/month — $90 more for the Opus premium. With 90% cache hit rate on Sonnet 4.6: approximately $11.40 + $75 = $86.40/month. For a startup processing customer queries, this is a practical monthly budget difference that compounds at scale.

Use-case fit

Candidate for: Production chat and assistant applications, code review, moderate-complexity generation, long-document workflows, and deployments where its lower listed rate may matter.

Compare alternatives if: The workload has costly failure modes, strict latency targets, or simple high-volume tasks. Test Opus, Haiku, and other eligible models on the same acceptance criteria rather than inferring fit from tier names.

Decision checklist

Before using Opus 4.8 instead: run a 50-sample eval of your hardest task category on both Sonnet 4.6 and Opus 4.8. If you can't detect a quality difference, stay on Sonnet — the 40% cost reduction compounds at production volume.

Before using Haiku 4.5 instead, compare it with Sonnet on the same tasks. Anthropic's published SWE-bench figures differ by about six points, but that benchmark gap does not establish the quality difference for classification, extraction, or your analysis workflow.

Frequently asked

How does Claude Sonnet 4.6 compare to Claude Opus 4.8 on cost?

Sonnet 4.6 is 40% lower on both listed input ($3 vs $5) and output ($15 vs $25) rates. Both models list a 1M-token context window. Anthropic reports 79.6% SWE-bench for Sonnet and 88.6% for Opus, but that benchmark gap does not translate directly to production quality. Compare them on the same workload and acceptance criteria.

What is the 64K max output limit for Claude Sonnet 4.6?

Sonnet 4.6 supports up to 64,000 output tokens per call. Opus 4.8 supports up to 128,000, so Sonnet does not have an output-length advantage over Opus. Choose between them using your required output length, evaluation results, and the $15 vs $25 per-million output rates.

Is Claude Sonnet 4.6 better than GPT-5 for most applications?

There is no provider-independent basis for a universal answer. The listed input rates are $3/1M for Sonnet 4.6 and $1.25/1M for GPT-5, while their published benchmark figures come from provider materials and may use different conditions. Run both on the same versioned tasks and compare accepted outputs, latency, retries, and total cost.

Changelog

  • — Replaced universal-default and production-quality claims with provider attribution and matched-workload evaluation guidance; synchronized metadata and FAQ copy.
  • — Corrected Sonnet 4.6 to a 1M-token context window and Opus 4.8 to a 128K max output.
  • — Expanded with Sonnet vs Opus comparison, 64K output analysis, caching economics, and cost scenarios.
  • — Published. Pricing verified at anthropic.com/pricing.

Sources