What this guide covers
This guide covers three hosted frontier candidates in 2026: Claude Opus 4.7, GPT-5, and Gemini 3.1 Pro preview. The linked reviews separate provider-published specifications and benchmarks from editorial hypotheses. The comparison supplies seven workload designs you can reproduce; it does not claim an unpublished benchr test or a universal winner.
Reviews
-
Claude Opus 4.7, reviewed
Anthropic positions Opus 4.7 for demanding coding and agentic work and publishes its price and benchmark sheet. Treat architectural fit as an editorial hypothesis: test it on held-out repository issues and document tasks, then compare errors and review time.
-
GPT-5, reviewed
OpenAI publishes GPT-5 specifications and benchmark results that make it a candidate for structured output and reasoning workloads. Speed, prose preference, and technical error rate depend on the deployment and prompt; measure them on the same local rubric.
-
Gemini 3.1 Pro, reviewed
Google documents the model's context capacity and Workspace integrations. Image-heavy and Workspace-bound work are reasonable editorial hypotheses to test with representative screenshots, documents, citations, and failure cases.
Comparisons
-
GPT-5 vs Claude Opus 4.7: a seven-task evaluation plan
Seven workload categories with a shared prompt and scoring template. The category leans are editorial hypotheses, not private measured results; reproduce them with pinned versions, retained outputs, and blind review.
-
Multimodal evaluation plan: twelve images, four models
A reproducible rubric for Claude, GPT-5, Gemini 3, and Llama 4 across dense interfaces, document images, charts, and Arabic script. Placements are hypotheses to validate, not an unpublished winner table.
-
The price-per-use-case table
Six workload shapes using dated, provider-listed token prices. Recalculate with your input/output mix, cache hit rate, tool calls, retries, and current price pages before making a budget decision.
Which one should you use?
Start with the candidate whose provider-documented capabilities match your dominant workload: Claude Opus 4.7 for a technical-work evaluation, GPT-5 for a structured-output evaluation, or Gemini 3.1 Pro preview for a multimodal or Workspace evaluation. These are editorial starting hypotheses, not measured winners.
To decide whether two subscriptions add value, route the same held-out set through each model alone and through the proposed pair. Record incremental task coverage, review time, failures, and the actual monthly bill; keep the second model only if the measured gain clears your threshold.
For screenshots, PDFs, or document images, include Gemini in the candidate set because Google documents native multimodal support. Compare extraction accuracy, spatial grounding, Arabic-script handling, citations, latency, and current provider-listed cost on your own corpus before expanding the stack.
For deeper context: the comparison tool lets you pick any of these models and any dimension to compare, with a downloadable PDF. The cost guide covers pricing dynamics in detail.