Two models, one budget, completely different machines. This is the rare comparison where you can ignore the rate card — a few cents per million tokens separates them in either direction depending on your input/output ratio. Everything that matters here is in the engineering columns.
Side-by-side specs
| Dimension | GPT-5 | Gemini 3.5 Flash |
|---|---|---|
| Input / 1M | $1.25 | $1.50 |
| Output / 1M | $10.00 | $9.00 |
| Cached input / 1M | — | $0.15 |
| Free API tier | No | Yes |
| Context window | 400,000 | 1,048,576 |
| Max output | 128,000 | 65,536 |
| First token (benchr est.) | 520ms | 195ms |
| Throughput (benchr est.) | 90 tok/s | 289 tok/s |
Speed is Flash's whole pitch
At 195ms to first token and 289 tokens/second in benchr's tracking, Flash is one of the fastest frontier-tier models on the market — about 3× GPT-5's throughput. For user-facing chat, live coding assistants, and any product where someone watches the answer stream, that difference is visible to the naked eye. A 2,000-token answer takes roughly 7 seconds to stream on Flash and 22 on GPT-5. No price discount buys back fifteen seconds of a user's attention.
Where GPT-5 holds the line
Output ceiling and ecosystem. GPT-5's 128K max output doubles Flash's 64K — relevant for long code generations and document drafting in one call. And the integration story repeats from every OpenAI comparison: more frameworks, more battle-tested function calling, more drop-in compatibility. On coding, benchr tracks Flash at 80.6% SWE-bench Verified against GPT-5's official 74.9% — read that gap as directional, since the Flash figure is an editorial estimate in benchr's index, flagged as such, while GPT-5's is official. The Flash review has the fuller picture.
The free-tier wrinkle
Flash has a free API tier; GPT-5 doesn't. For prototyping, internal tools, and low-volume side projects, that's not a rounding error — it's the entire bill. Plenty of teams should prototype on Flash's free tier, measure, and only then decide where paid traffic goes. The free-AI roundup maps where the free tier taps out.