Side-by-side specs
| Dimension | GPT-5 | Claude Opus 4.8 |
|---|---|---|
| Provider | OpenAI | Anthropic |
| Released | Aug 7, 2025 | May 28, 2026 |
| Input / 1M | $1.25 | $5.00 |
| Output / 1M | $10.00 | $25.00 |
| Cached input / 1M | — | $0.50 (90% off) |
| Context window | 400,000 tokens | 1,000,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| SWE-bench Verified | 74.9% | 88.6% |
| GPQA Diamond | not published | 93.6% |
| Fast mode | No | Yes — $10/$50, ~2.5× speed |
The price gap in real numbers
At 10 million input tokens per month — a moderate agentic workload — you're spending $12.50 on GPT-5 or $50 on Opus 4.8. That $37.50/month difference is $450/year. At production scale of 500M input tokens monthly, it's $7,500 vs $30,000.
Opus's caching changes the math. Claude caches system prompts at $0.50/1M — a 90% discount. If your agent sends a 40,000-token context on every call, the effective input cost per call drops sharply once the cache is warm. GPT-5 has no published caching discount. For agentic workloads with large, static system prompts, Opus's all-in monthly cost can be closer to GPT-5's than the per-token headline rates suggest. Model your actual call patterns before assuming GPT-5 is cheaper.
What the 14-point coding gap actually means
SWE-bench Verified presents real GitHub issues from open-source Python repositories. The model must read the issue, understand the code, write a fix, and pass the existing tests — no scaffolding, no hints. At 88.6%, Opus resolves roughly 88 of 100 issues without human intervention. GPT-5 resolves about 75.
That gap matters most when humans are out of the loop. If you're running a CI bot that autonomously applies model-generated patches, those 13 additional failures per 100 issues translate directly to broken builds or manual cleanup. At 200 automated PRs per day, that's 26 more failures requiring engineer attention. At $30 per engineer-hour, the math moves fast.
Outside fully automated code repair, the gap narrows considerably. Writing functions from spec, drafting docstrings, explaining unfamiliar code — GPT-5 handles all of these well enough that most teams won't notice the difference. The premium is for removing humans from the loop, not assisting them.
Context window: when 400K isn't enough
For most production workloads, GPT-5's 400K-token context is sufficient. Where it bites: entire codebases, long legal contracts, multi-session agent memory, or document analysis jobs that want to fit everything in at once. Opus 4.8's 1M-token context handles all of these without truncation. If your use case regularly approaches 300K tokens, this is a functional constraint, not a spec comparison.