Claude Opus 4.7, reviewed

Coding, long-document analysis, and multilingual capability. What Opus 4.7's pricing and documented capability profile imply for the workloads you should put on it.

By benchr Editorial Team · · View changelog · Provider figures are dated snapshots; methodology reviewed July 23, 2026

benchr rating: 4.6 / 5

Claude Opus 4.7, reviewed: warm clay layers and branching decision routes.
Benchr editorial field plate Claude Opus 4.7 Layered reasoning · Anthropic family
AnthropicClaude Opus 4.7, reviewed is mapped with warm clay layers and branching decision routes.
Input cost / 1M $5.00 Anthropic list-price snapshot; verify live pricing
SWE-Bench Verified 87.6% Anthropic-reported on SWE-bench Verified; not a production guarantee
Context window 1M Maximum advertised by Anthropic; effective quality must be tested
Output cost / 1M $25 Anthropic list-price snapshot; Sonnet 4.6 was $15

Anthropic announced Claude Opus 4.7 on April 16, 2026, per its launch announcement. This review treats the launch chart, model documentation, and pricing as provider evidence. Those sources describe the product and its test configuration; they do not establish how it will perform or what it will cost in a particular production system.

This piece reads the public evidence: the launch announcement, the published Claude API documentation, the SWE-bench Verified score Anthropic reports, and the pricing page. It turns those claims into evaluation hypotheses. It does not supply an independent benchr performance result or a universal workload assignment.

If your workload involves production code, dense documents, or costly errors, include Opus in a controlled comparison. Include a lower-cost tier as the baseline, keep prompts and tool access fixed, and decide only after measuring task success, reviewer effort, latency, and token use.

What the documented capability profile says

Anthropic positions Opus 4.7 around coding, reasoning under uncertainty, and long-document analysis. Its materials cite a provider-reported SWE-bench Verified result, describe multi-step problem-solving, and advertise a 1M-token context window. These are useful inputs for selecting test cases, but neither a benchmark score nor a context maximum guarantees production task completion.

Weigh these as vendor capability claims. They indicate intended use and test emphasis, not independent verification. Test any candidate—including Opus—on your own inputs, tooling, safety constraints, and acceptance criteria before drawing a capability or value conclusion.

Coding: test whether the price difference earns its keep

Anthropic reports Opus 4.7 at 87.6% on SWE-bench Verified in the cited release material. Treat that as a provider-reported result under a specific benchmark harness, not a success rate for your repository and not proof of superiority over another model or configuration.

A reproducible repository test should sample real issue types: a bounded bug fix, a cross-file feature, an architectural proposal, and an unfamiliar-library change. Give each candidate the same repository state, prompt, tools, time limit, and retry budget. Record test-pass rate, unauthorized edits, fabricated APIs, reviewer minutes, latency, and total tokens. That evidence can show whether Opus's premium is justified for your codebase.

For a mechanical change such as renaming a variable across 20 files, test Sonnet and Opus against the same acceptance suite. The lower list price makes Sonnet a reasonable baseline, but price alone does not show that it will complete the task correctly.

HARD WORK ↑ COST SENSITIVE → Opus 4.7 production code · long docs Sonnet 4.6 daily driver · drafts · routine code Haiku 4.5 high-volume · classification
Anthropic's product positioning for three Claude tiers, captured in May 2026. Treat the placements as provider guidance and verify the live catalog and task fit before deployment.

Long documents: test retrieval and summarization separately

Anthropic advertises a 1M-token context maximum for Opus 4.7. Capacity is not a retrieval, citation, or summarization-quality guarantee. Effective limits depend on document structure, fact position, prompt design, tool use, and the error threshold of the application.

To test a long-document workflow, create documents with answer keys and place relevant facts near the beginning, middle, and end. Compare a one-pass summary with a query-then-summary workflow under the same token budget. Score cited-fact accuracy, omissions, unsupported statements, and reviewer time. Adopt the two-pass design only if it improves those measures on your documents. The context-window guide explains why the advertised maximum and effective task limit are different quantities.

A context maximum is a capacity specification. Retrieval and summarization quality still require separate tests.

For long-document work such as legal review or policy analysis, compare models at the actual document lengths and risk level you expect. A smaller model may require chunking or retrieval, while a larger window may increase direct capacity; neither architecture establishes accuracy without an answer-key evaluation.

Multilingual: test dialect and register directly

Public multilingual benchmarks do not establish dialect or register fit for a specific audience. Treat English, Gulf Arabic, romanized Hindi, and Indonesian-market copy as separate test groups; one multilingual score cannot stand in for all four.

Build a blind review with native speakers, fixed prompts, and a rubric for factual fidelity, grammar, dialect, formality, and editing time. Use enough examples from each target market to expose variation, and publish the sample and rubric with any conclusion. For an Arabic-specific protocol, see benchr's Arabic-content evaluation guide.

Failure hypotheses to test before deployment

Community anecdotes are not a measured failure rate. The following are test hypotheses for a pre-deployment suite, not claims that Opus exhibits them at a known frequency.

First hypothesis: output-length control. Run a fixed set of yes-or-no questions with and without a one-line constraint. Measure instruction compliance, factual accuracy, and output tokens; do not assume one prompt amendment will work across calls.

Second hypothesis: fabricated API signatures. Sample both standard and less-common libraries, pin each library version, and check every generated symbol against official documentation or a compiler. Report fabrication and abstention rates by library group; do not presume reliability for either group.

Third hypothesis: scope drift during refactors. Define an allowed-file and allowed-symbol list, run identical tasks, and count out-of-scope edits. Compare an explicit scope constraint with the baseline prompt before choosing a production guardrail.

What it costs

In the pricing snapshot used for this article, Anthropic listed Opus 4.7 at $5 per million input tokens and $25 per million output tokens, with lower cached-input rates. These are provider-published list prices, not a quote or a prediction of total production cost. Verify the live pricing page and include retries, tool calls, cache behavior, and review labor in your calculation.

List pricing reference, May 2026 — Anthropic Claude family
TierInput / 1MOutput / 1MWorkload fit
Opus 4.7$5.00$25Production code, long-document work, hard reasoning
Sonnet 4.6$3.00$15Daily driver: drafts, chat, routine code
Haiku 4.5$1.00$5High-volume, classification, extraction

Do not infer a default tier or escalation rule from list prices alone. For each workload class, measure pass rate, human-review minutes, latency, token volume, retries, and error impact. A lower-cost tier is economical only if it stays within the acceptance threshold; an Opus escalation is justified only if its measured benefit exceeds its incremental total cost.

1.67× Opus 4.7's input price vs Sonnet 4.6. Output gap is wider.

A comparison with GPT-5 or another Claude tier should use the same prompts, tools, limits, and scoring rubric. The head-to-head guide provides a reproducible evaluation design. A multi-model router is an engineering option, not an automatic saving: test its classification errors, fallback rate, latency, and total token cost before deployment. The price-per-use-case guide shows how to model those inputs.

Evidence-based conclusion

The public evidence supports testing Claude Opus 4.7 for coding, long-document analysis, and reasoning workloads; it does not make the model a universal default. Anthropic's published prices and benchmark results define a comparison starting point. Your acceptance tests determine whether any measured quality or review-time gain justifies the price difference.

Before choosing one tier or a two-tier router, run a shadow evaluation on representative traffic and predefine its pass, cost, and latency thresholds. Keep the simpler configuration if routing does not produce a measured net benefit after classification mistakes, retries, and operating overhead.

Frequently asked

Is Claude Opus 4.7 worth $5 per million input tokens?

There is no universal yes-or-no answer. Anthropic listed Opus 4.7 at $5 per million input tokens and reported an 87.6% SWE-bench Verified score in this article's source snapshot, but that benchmark does not establish return on investment for your workload. Run the same representative tasks on Opus and Sonnet, blind-review the outputs, and pay the premium only if the measured quality or error-cost difference justifies it.

How does Claude Opus 4.7 compare to GPT-5?

Published scores and provider positioning are not enough to name a universal winner. Compare both models on the same versioned tasks and prompts, then record task success, factual or code-review errors, latency, total tokens, and cost. Provider-reported benchmarks are starting evidence, not a production guarantee.

What does Claude Opus 4.7 cost?

In the provider-pricing snapshot used for this article, Anthropic listed $5 per million input tokens and $25 per million output tokens, with lower cached-input rates. Pricing can change, so verify the live provider page and calculate spend from your measured input, output, and cache volumes.

When should you use Sonnet 4.6 instead of Opus 4.7?

Treat Sonnet as a lower-cost candidate for repeatable, lower-risk work, but do not assume it is sufficient by task label alone. Build a representative evaluation set, define acceptable error and review-time thresholds, and choose the least-expensive tier that passes those thresholds.

What's the effective context window for Claude Opus 4.7?

Anthropic advertised a 1M-token maximum context window. That capacity does not guarantee retrieval or summarization quality. Test documents at several lengths and fact positions, then score citation accuracy, missed facts, unsupported claims, and summary omissions before setting a production limit.

Changelog

  • July 23, 2026 — Replaced anecdotal behavior, universal tier, routing, and return-on-investment claims with reproducible evaluation protocols. Labeled provider-reported benchmarks, context, and pricing as dated inputs rather than production guarantees, and synchronized the FAQ schema.
  • May 25, 2026 — Pre-publication draft: rewrote sections that previously narrated original lab tests; the article now distinguishes published benchmarks, pricing, and Anthropic's own positioning.
  • May 4, 2026 — Retrospective coverage note incorporated at publication: a Sonnet 4.6 comparison was added for the cross-tier math.
  • May 30, 2026 — Published with retrospective coverage through the April 22, 2026 subject date.

References

  1. Anthropic, "Claude API Documentation," docs.claude.com, accessed May 2026.
  2. Anthropic, "Claude Pricing," anthropic.com/pricing, accessed May 2026.
  3. LMSYS, "Chatbot Arena leaderboard," lmarena.ai, May 2026 snapshot.
  4. "SWE-bench Verified leaderboard," swebench.com, May 2026.
  5. Anthropic, "Introducing Claude Opus 4.7," anthropic.com/news/claude-opus-4-7, April 16, 2026.