Both models sit below their providers' higher-priced tiers and are positioned for sustained API use. The listed prices, context limits, and reported benchmark results differ, but none of those fields shows production reliability by itself. Compare them on the same prompts, tool calls, and acceptance criteria.
Side-by-side specs
| Dimension | Claude Sonnet 4.6 | GPT-5 |
|---|---|---|
| Input / 1M | $3.00 | $1.25 |
| Output / 1M | $15.00 | $10.00 |
| Cached input / 1M | $0.30 | $0.125 |
| Context window | 1,000,000 | 400,000 |
| Max output | 64,000 | 128,000 |
| SWE-bench Verified | 79.6% | 74.9% |
| Throughput (benchr est.) | 95 tok/s | 90 tok/s |
GPT-5 stays cheaper when you cache
Both models discount cached input by 90%: Sonnet charges $0.30/1M and GPT-5 charges $0.125/1M. For an agent with a reusable 40K-token system prompt, that prefix costs $0.012 per Sonnet call versus $0.005 per GPT-5 call once it is cached. Stateless, short-prompt workloads stay on the standard $3 and $1.25 input rates instead. Run your own mix through the cost calculator.
A concrete workload
Take a code-review bot: 15K input tokens (diff + context), 3K output, and 20,000 runs a month. Without caching, GPT-5 costs $975/month and Sonnet 4.6 costs $1,800/month. If 10K input tokens are a stable cached prefix, the totals fall to roughly $750/month on GPT-5 and $1,260/month on Sonnet.
The remaining difference is about $510 a month for a 4.7-point gap in the cited SWE-bench figures. That benchmark alone cannot tell you whether the premium reduces missed bugs in your review queue, so measure accepted reviews and developer corrections.
Workloads to test for each model
Sonnet 4.6 has the higher cited SWE-bench figure and a context window above 400K. GPT-5 has the lower standard and cached input price, the larger listed output limit (128K vs 64K), and may require less integration work in a stack already using OpenAI API shapes. Choose against your token mix and representative repository tasks; the Sonnet review provides broader context.