Gemini 3.1 Pro
The >200K-token tier applies to batch pricing as well: batch input doubles from $1 to $2 per 1M and batch output rises from $6 to $9.
The published record
Every number below was read from the provider's own documentation on the date shown. benchr does not restate a figure it has not seen published.
- API identifier
gemini-3.1-pro-preview- Context window
- 1MtokensSource
- Maximum output
- 64KtokensSource
- Input
- $2.00per 1M tokensSource
- Output
- $12.00per 1M tokensSource
- Cached input
- $0.2per 1M tokensSource
- Released
- February 19, 2026Source
- License
- Proprietary
Availability Preview (GA coming soon).
Record verified May 31, 2026 — read 98 days ago, and a price can move in a week
What the record says
Tiered pricing: the >200K-token tier doubles input ($4) and raises output ($18). Output price includes thinking tokens. No free API tier (free trial in AI Studio UI only). Context-caching: cached input $0.20/1M (90% off the $2 input rate; Google additionally charges $1 per 1M-token-hour of cache storage), verified 2026-06-15 against ai.google.dev/gemini-api/docs/pricing.
What benchr has documented it doing
Capabilities in the benchr ledger that name this model. Documented means a provider says it works and benchr recorded where. Verified means benchr ran it.
2 capabilities name this model. All 2 are partly documented by the provider; benchr has run 0 of them itself.
Which one to use
The rest of the family, with the two numbers that usually decide it. The current model is marked.
| Models | Input | Output | Context window |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | $0.3 | $2.50 | 1,048,576 |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1,048,576 |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1,048,576 |
| Gemini 3.1 Pro This page | $2.00 | $12.00 | 1M |
benchr's read
Worth it for
- Deep reasoning in the Gemini family
- Long-context vision work
- Workspace integration
Look elsewhere if
- Coding agents — Flash is faster and cheaper
- Cost-sensitive workloads — note the over-200K price bump
Prices, limits, and identifiers above are provider facts. This section is benchr's judgement about them.
What benchr has written about it
Pieces that name this model, newest first. The ones written about this model come before the ones that mention it in passing.
What is not on this page
Stated rather than filled in.
- No first-token or tokens-per-second figure is recorded. benchr has not measured it and the provider does not publish one.
- benchr has not run any capability on this model itself. Everything in the capability section is documentation, not a test result.