This is the daily-driver decision most teams face in 2026. Neither model is the flagship of its lab — Opus 4.8 and GPT-5.5 sit above them — but these two carry the bulk of real production traffic because they're the tiers priced for volume. They're built to different recipes, and the price sheet only tells half of it.
Side-by-side specs
| Dimension | Claude Sonnet 4.6 | GPT-5 |
|---|---|---|
| Input / 1M | $3.00 | $1.25 |
| Output / 1M | $15.00 | $10.00 |
| Cached input / 1M | $0.30 | — |
| Context window | 1,000,000 | 400,000 |
| Max output | 64,000 | 128,000 |
| SWE-bench Verified | 79.6% | 74.9% |
| Throughput (benchr est.) | 95 tok/s | 90 tok/s |
The price gap shrinks when you cache
The headline says Sonnet costs 2.4× more on input. The footnote says Sonnet has a $0.30 cached-input tier and GPT-5 doesn't have a cached rate in benchr's verified record. For an agent with a 40K-token system prompt called thousands of times a day, that changes the ranking: the repeated prompt bills at $0.30/1M on Sonnet, under a quarter of GPT-5's $1.25 standard rate. Stateless, short-prompt workloads never see this benefit — which is exactly why the right answer differs per team. Run your own mix through the cost calculator.
A concrete workload
Take a code-review bot: 15K input tokens (diff + context), 3K output, 20,000 runs a month. On GPT-5 that's $375 input + $600 output = $975/month. On Sonnet 4.6 without caching: $900 + $900 = $1,800/month. If 10K of those input tokens are a stable cached prefix, Sonnet drops to roughly $1,260/month. You're paying a few hundred dollars a month for a 4.7-point SWE-bench edge — cheap if it cuts even one missed bug per week, pointless if your reviews are simple style checks.
Where each one wins
Sonnet 4.6 wins on repository-scale coding, anything that needs more than 400K tokens in one window, and prompt-heavy agents that exploit caching. It's the model the benchr review calls the daily-driver tier for a reason. GPT-5 wins on raw price for short tasks, on output ceiling (128K vs 64K max output), and when your stack is already built on OpenAI's API shapes. The wrong reason to pick either is brand loyalty — the right reason is your token mix.