Gemini 3.7 Flash
Google · mid-tier. Coding and agent work at Flash prices.
The published record
Every number below was read from the provider's own documentation on the date shown. benchr does not restate a figure it has not seen published.
- API identifier
gemini-3.7-flash- Context window
- 1,048,576tokensSource
- Maximum output
- 65,536tokensSource
- Input
- $0.75per 1M tokensSource
- Output
- $3.75per 1M tokensSource
- Cached input
- $0.075per 1M tokensSource
- Released
- August 13, 2026Source
- License
- Proprietary
Availability GA (stable) since August 13, 2026.
Record verified August 24, 2026
What the record says
Google made Gemini 3.7 Flash generally available on August 13, 2026 and describes it as its most intelligent workhorse model for coding and agents, with improvements across software engineering, web development, and agentic workflows. The model page lists text, image, video, audio, and PDF input with text output, a 1,048,576-token input limit, a 65,536-token output limit, caching, code execution, file search, function calling, Search and Maps grounding, structured outputs, thinking at low, medium, or high (no minimal level), URL context, and Computer Use in preview. Live API, image generation, and audio generation are not supported. Google publishes no benchmark figures for it, so none are claimed here.
What benchr has documented it doing
Capabilities in the benchr ledger that name this model. Documented means a provider says it works and benchr recorded where. Verified means benchr ran it.
2 capabilities name this model. All 2 are officially supported by the provider; benchr has run 0 of them itself.
What has moved
Entries in the change ledger that affect this model, newest first.
Released
Google released gemini-3.7-flash as stable GA on 2026-08-13. Introductory Standard pricing through 2026-12-31 is $0.75/$3.75 per 1M with $0.075 cached input; scheduled 2027 Standard pricing is $1.50/$7.50. Source: ai.google.dev Gemini API changelog, model, latest-model, and pricing pages (verified 2026-08-21).
Which one to use
The rest of the family, with the two numbers that usually decide it. The current model is marked.
| Models | Input | Output | Context window |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | $0.3 | $2.50 | 1,048,576 |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.7 Flash This page | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1,048,576 |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M |
benchr's read
Worth it for
- Coding and agent work at Flash prices
- Teams already on Gemini 3.6 Flash — same price, newer model
- Long-context multimodal input with text output
- Free-tier prototyping before a paid rollout
Look elsewhere if
- You need published official benchmark tables — Google shipped none
- You want the cheapest subagent tier — use Gemini 3.5 Flash-Lite
- You need image generation, audio output, or the Live API
Prices, limits, and identifiers above are provider facts. This section is benchr's judgement about them.
What benchr has written about it
Pieces that name this model, newest first. The ones written about this model come before the ones that mention it in passing.
What is not on this page
Stated rather than filled in.
- No first-token or tokens-per-second figure is recorded. benchr has not measured it and the provider does not publish one.
- benchr has not run any capability on this model itself. Everything in the capability section is documentation, not a test result.