By benchr Editorial Team · Published September 7, 2026Provider-published facts rechecked against the official sources on September 7, 2026

Gemini 3.8 Flash: same rate card, different bill

Google's newest Flash endpoint carries the same published prices as the one it follows. The difference is how many tokens it decides to spend.

Two matching price cards feed reasoning chains of different length, one short and one long.
Benchr model field plate Gemini 3.8 Flash Same rates · longer runs
GoogleGemini 3.8 Flash: same rate card, different bill is framed by spectrum bands and long-context tracks.
Input / 1M$0.75Output: $3.75
Context1.05M65,536 max output
Price changes1 Jan 2027Every tier doubles
Released2 Sep 2026Twenty days after 3.7

Google made Gemini 3.8 Flash generally available on September 2, 2026, twenty days after Gemini 3.7 Flash. benchr rechecked the pricing table, the model page and the launch post on September 7. The comparison that matters here is unusual, because the two endpoints are not separated by price, context, or output limit. They are separated by behavior, and behavior is the part no rate card prices.

The rate card is the same one, to the cent

Every published lane matches its 3.7 Flash equivalent exactly. This is not an approximation from a summary table; it is what both pricing entries state.

Published rates per 1M tokens, Gemini 3.7 Flash and 3.8 Flash
LaneThrough 31 Dec 2026From 1 Jan 2027Differs from 3.7?
Standard input$0.75$1.50No
Standard output$3.75$7.50No
Cached input$0.075$0.15No
Cache storage / 1M / hour$0.50$1.00No
Batch and Flex input$0.375$0.75No
Batch and Flex output$1.875$3.75No
Priority input$1.35$2.70No
Priority output$6.75$13.50No

The input window is 1,048,576 tokens and the maximum output is 65,536 on both. A free tier is offered on both. So a migration cannot be justified or rejected on the price sheet, and a cost estimate that only swaps the model ID will return the same number for two models that do not behave the same way.

What Google changed is how much the model spends

The launch post is direct about it. Google writes that 3.8 Flash works harder, and that on complex tasks it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. Thinking is exposed at low, medium and high levels on the model page.

Read as a cost statement rather than a quality one, that sentence says the model will, by design, produce more output tokens and more tool calls for the same prompt. Output is priced at five times input on this rate card, and thinking tokens are billed as output. A workload that was comfortable on 3.7 Flash can therefore cost materially more on 3.8 Flash at identical published rates, and the increase will not appear anywhere except the token counter.

Google says as much in its own guidance: for applications where compute efficiency is the primary constraint, developers can use lower effort levels or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads. That is a vendor telling you the newer model is not automatically the right one.

One published number, and the ones that are not numbers

Google publishes a single figure for the general model at launch: 54.9% on HLE-Verified. That is a provider-reported result, not a benchr measurement, and benchr has not reproduced it.

The rest of the launch claims are comparative statements without published figures — that 3.8 Flash outperforms most larger frontier models on DeepSWE v1.1, and beats 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. Those may well be true. They are not usable in a procurement comparison, because a claim with no number cannot be checked, ranked, or entered into a spreadsheet. benchr records the one figure and leaves the others as what they are.

The Cyber variant is not something you can plan around

The same announcement introduces Gemini 3.8 Flash Cyber, a specialized variant for vulnerability detection and automated patching, with 47.2% pass@1 on CWE-Bench and a stated real-world vulnerability discovery rate above 70%.

It is reachable only through Google's Fairwind Program, which is restricted to trusted government authorities, critical infrastructure operators and software maintainers. It has no entry in the API pricing documentation and no separate API model ID published there, so benchr carries no record for it. If you are not already in that program, treat it as unavailable rather than as a model you might adopt later.

What the endpoint does and does not do

The model accepts text, image, video, audio and PDF input and returns text. Function calling, structured outputs, code execution, file search, URL context, Search grounding, Grounding with Google Maps, caching, Batch, Flex and Priority inference are all listed, and computer use is listed as Preview.

Image generation, audio generation and the Live API are explicitly not supported for this model. If a stage of your pipeline produces images or audio, or holds a live session, that stage belongs on a different endpoint regardless of how the reasoning stage performs.

Choosing between 3.7 and 3.8

Which Flash endpoint the official record points at
If your constraint isThe record points atWhy
Cost per completed task, on work the current model already handlesGemini 3.7 FlashGoogle names it as the efficiency-first choice and keeps it fully supported
Multi-step agentic work that currently stalls or gives up earlyGemini 3.8 FlashExtra reasoning steps and iterative tool calls are the stated change
A fixed per-request token budgetGemini 3.7 Flash, or 3.8 at a low thinking levelHigher effort levels spend output tokens, which are billed at five times input
Generated images, generated audio, or a live sessionNeitherThose output paths are not listed for either model
A budget that crosses into 2027Model both periods firstEvery tier doubles on January 1, 2027 on both models

What a pilot has to measure

Because the prices are identical, a token-rate comparison tells you nothing. The measurement that separates these two models is cost per completed task, not cost per token, and it needs the same representative cases run through both endpoints with the tool calls counted.

Record, per case: total input tokens, total output tokens including thinking, the number of tool calls, whether the task completed without human intervention, and the wall-clock time. A model that costs 40% more per task and completes 20% more of them is a different decision from one that costs 40% more and completes the same set. Neither of those outcomes is visible on the rate card, and neither is predictable from the launch post.

If a workload uses repeated prefixes, price context caching properly on both: the cache-hit rate is a per-model behavior, and the estimate has to carry cache storage at $0.50 per million tokens per hour rather than treating every hit as free.

Frequently asked

What is the Gemini 3.8 Flash API model ID?

Google lists gemini-3.8-flash as the model ID.

Is Gemini 3.8 Flash more expensive than 3.7 Flash?

Not on the published rate card — every lane is priced identically through December 31, 2026, and every lane doubles on January 1, 2027 on both models. Google states that 3.8 Flash executes extra reasoning steps and calls tools iteratively, so the same task can consume more tokens and cost more in practice.

Should I migrate from Gemini 3.7 Flash?

Google keeps 3.7 Flash fully supported and names it for efficiency-first workloads. Migrate when your work involves multi-step agentic tasks that benefit from more reasoning steps, and measure cost per completed task rather than cost per token before committing.

Can I use Gemini 3.8 Flash Cyber?

Only through Google's Fairwind Program, which is limited to trusted government authorities, critical infrastructure operators and software maintainers. It is not published in the API pricing documentation.

Changelog

  • September 7, 2026 — Published after rechecking the official pricing table, model page and launch announcement, and after adding the one provider-published benchmark figure to the model record.

References

  1. Official release announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
  2. Official pricing documentation: https://ai.google.dev/gemini-api/docs/pricing
  3. Official model documentation: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash
  4. Official provider documentation: https://ai.google.dev/gemini-api/docs/deprecations