GPT-5 vs Gemini 3.5 Flash: price, context, and latency

A few cents separate the rate cards. The practical differences are context, output limits, cache pricing, features, and the free tier; latency still needs a controlled deployment test.

By benchr Editorial Team · · View changelog · Provider fields rechecked; undocumented benchmark and latency estimates removed

GPT-5 input / 1MOpenAI · output $10
Flash input / 1MGoogle · output $9
Flash max outputGPT-5: 128K
Flash contextGPT-5: 400K

Two models, one budget, completely different machines. This is the rare comparison where you can ignore the rate card — a few cents per million tokens separates them in either direction depending on your input/output ratio. Everything that matters here is in the engineering columns.

Side-by-side specs

GPT-5 vs Gemini 3.5 Flash — provider-published price and limit fields
DimensionGPT-5Gemini 3.5 Flash
Input / 1M$1.25$1.50
Output / 1M$10.00$9.00
Cached input / 1M$0.125$0.15
Free API tierNoYes
Context window400,0001,048,576
Max output128,00065,536

Latency needs a deployment test

Neither provider publishes one universal, directly comparable tokens-per-second or time-to-first-token guarantee for these two models, and benchr has not run a controlled cross-region study. The undocumented values formerly shown here were removed. Benchmark the same prompts, output lengths, settings, and concurrency in your target deployment and compare p50, p95, and p99 latency.

Reasons to evaluate GPT-5

GPT-5's listed 128K max output is larger than Flash's listed 65,536-token cap, which can matter for long code generations or documents. For coding, OpenAI publishes a SWE-bench figure for GPT-5, while the checked Google material does not publish a directly comparable SWE-bench Verified figure for Gemini 3.5 Flash. Run both on the same repository tasks and settings rather than filling the gap with an estimate.

Free-tier availability

Flash has a free API tier; GPT-5 doesn't. For prototyping, internal tools, and low-volume side projects, that's not a rounding error — it's the entire bill. Plenty of teams should prototype on Flash's free tier, measure, and only then decide where paid traffic goes. The free-AI roundup maps where the free tier taps out.

Frequently asked

GPT-5 and Gemini 3.5 Flash cost almost the same — what's the real difference?

The published context and output limits, features, and ecosystem fit differ. Neither provider supplies one directly comparable universal latency figure for this page, so benchmark both models in your deployment rather than assuming a fixed multiple.

Which is cheaper for a chat application?

Calculate with your actual input/output mix, cache hit rate, and current provider terms. A single example workload cannot establish the cheaper model for every chat application.

Does Gemini 3.5 Flash beat GPT-5 on coding?

This page cannot establish that. OpenAI publishes a SWE-bench figure for GPT-5, while the checked Google material does not publish a directly comparable SWE-bench Verified figure for Gemini 3.5 Flash. Test both on the same repository tasks and settings.

Changelog

  • July 30, 2026 — Rechecked both provider records and removed undocumented latency and Gemini SWE-bench estimates.
  • July 24, 2026 — Corrected GPT-5 cached-input pricing to $0.125/1M.
  • June 10, 2026 — Expanded from a spec table into a full comparison with rate-card arithmetic, published limits, estimate caveats, and a decision checklist. Page re-indexed.
  • June 6, 2026 — Published as a spec-table stub (noindexed pending expansion).

Sources