Gemini 3.1 Pro API pricing: 1M context and 94.3% GPQA at $2/1M

Gemini 3.1 Pro leads on GPQA Diamond (94.3%) and offers the largest commercial context window at 1M tokens — enough for entire codebases, long documents, or dozens of research papers in a single call. At $2/1M input it sits in the mid-range price tier while delivering frontier science and reasoning benchmarks.

By the benchr team · · Figures verified against official sources, June 6, 2026 · View changelog

Input / 1MGoogle AI
Output / 1MGoogle AI
GPQA Diamondverified
Contextmax window

Pricing breakdown

gemini-3.1-pro — official Google AI pricing
TierRate / 1M tokens
Standard input (≤200K)$2.00
Standard input (>200K)$4.00
Standard output$12.00
Context caching$0.20
Context window1,000,000 tokens

The 1M context window — practical applications

One million tokens is approximately 750 pages of text, 150–200 images, or 2–3 hours of video transcript. This enables workflows that are impractical or impossible with smaller context models: loading an entire 50K-line codebase into context for architectural review; processing a full legal contract package; analyzing a multi-hour video recording in a single call; synthesizing findings across 50 research papers simultaneously. For any task where chunking introduces coherence loss, Gemini 3.1 Pro's 1M window removes the constraint.

Note the tiered pricing: inputs over 200K tokens cost $4/1M rather than $2/1M. For workloads that regularly use 500K–1M tokens of context, the effective input cost per call is higher — budget accordingly.

GPQA Diamond: 94.3% on PhD-level science

GPQA Diamond tests PhD-level questions in biology, chemistry, and physics — questions that require domain expertise beyond pattern matching. Gemini 3.1 Pro's 94.3% score leads the benchmark among major commercial models, ahead of Claude Opus 4.8 at 93.6% and GPT-5.5. For research assistance, scientific analysis, medical reasoning, and technical document understanding, this performance level is directly relevant to task quality. For general-purpose coding or business tasks, GPQA is less discriminating than SWE-bench.

Native multimodal: video, audio, images, PDF

Gemini 3.1 Pro processes text, images, video, audio, and PDF in the same API call. This is architecturally different from adding vision as an afterthought: Google trains multimodal understanding from the ground up. A workflow that requires analyzing a technical presentation (PDF + slides + audio) can be submitted as a single call. For document intelligence, media analysis, and cross-modal research tasks, Gemini 3.1 Pro is the strongest commercial offering at its price point.

Cost scenarios

At 10M input + 2M output per month within 200K context tiers: $20 + $24 = $44/month. Claude Opus 4.8 at the same volume: $50 + $50 = $100/month — Gemini is 56% cheaper. For long-context use (averaging 500K tokens per call): 10M total input at $4/1M blended = $40 + $24 output = $64/month — still 36% below Opus 4.8 at the same output volume.

Use-case fit

Best for: Long-document analysis requiring more than 1M tokens of context; scientific research and technical reasoning where GPQA leadership matters; multimodal pipelines combining text, video, audio, and images; Google Cloud/Vertex AI integrated stacks.

Skip if: Your primary need is SWE-bench coding performance — Gemini 3.1 Pro's coding benchmark trails Claude Opus 4.8 and GPT-5.5. Also skip for simple volume tasks where the $2/1M base rate is higher than necessary — Gemini 3.5 Flash at $1.50/1M handles lighter workloads.

Decision checklist

Identify your actual context length requirements: if your p90 context is under 200K tokens, the 1M window is not a differentiator. Gemini 3.5 Flash at $1.50/1M covers most tasks at lower cost. Use Gemini 3.1 Pro specifically when the context ceiling or GPQA performance gap is the bottleneck.

For multimodal use: confirm whether your input modalities (video, audio) actually benefit from the native architecture. For simple image captioning or document OCR, lighter models are sufficient.

Frequently asked

What can you do with Gemini 3.1 Pro's 1M context window?

Approximately 750 pages of text, entire mid-size codebases, 4–6 hours of video transcript, or hundreds of research papers in a single call. Practical: full codebase review, processing complete legal document sets, video analysis, multi-paper synthesis. No other commercial model matches this at $2/1M input pricing.

How does Gemini 3.1 Pro compare to Claude Opus 4.8 on benchmarks?

Gemini 3.1 Pro leads GPQA Diamond (94.3% vs 93.6%). Opus 4.8 leads SWE-bench (88.6%). Gemini 3.1 Pro costs $2/1M input vs Opus 4.8 at $5/1M — 60% cheaper. For science and research, Gemini 3.1 Pro is more cost-efficient. For coding pipelines, Opus 4.8 is stronger.

Does Gemini 3.1 Pro support video and audio input?

Yes — natively. Gemini 3.1 Pro handles text, images, video, audio, and PDF in the same call. This native multimodal architecture enables workflows combining multiple media types without separate processing steps.

Changelog

  • — Expanded with 1M context analysis, GPQA benchmarks, multimodal capabilities, and cost scenarios.
  • — Published. Pricing verified at ai.google.dev/pricing.

Sources

  • Google AI Studio pricing — ai.google.dev/pricing (verified June 6, 2026)
  • GPQA Diamond leaderboard — huggingface.co/spaces/opencompass (verified June 6, 2026)
  • benchr models.json — verified June 6, 2026