Gemini 3.7 Flash costs the same as 3.6 Flash — and both double in January

Google's newest Flash model shipped with a discount that has a hard expiry date. Two questions follow, and only one of them is about capability.

By benchr Editorial Team · · View changelog · Figures verified against Google's live Gemini API pricing and model pages, August 24, 2026

Gemini 3.7 Flash pricing plate: two Flash tiers at one rate, with a January 1 step in the price line.
Benchr model field plate Gemini 3.7 Flash Introductory rate · dated
GoogleGemini 3.7 Flash reached general availability on August 13, 2026 at a rate that expires December 31, 2026.
Input / 1M$0.75Through December 31, 2026
Output / 1M$3.75Includes thinking tokens
From Jan 1, 2027$7.50Output; input returns to $1.50
Context1.05M1,048,576 in; 65,536 out

Google shipped Gemini 3.7 Flash on August 13, 2026 and calls it the most intelligent workhorse model it has built for coding and agents. That's the pitch. The part you can verify without running a single request is the pricing table, and it's doing something unusual: the new model and the model it succeeds are listed at exactly the same numbers, down to the cache rate.

Two models, one price, one deadline

Here's the current Flash ladder as Google publishes it. The promotional column is what you pay today; the January column is what the same workload costs four months from now if you change nothing.

Google Gemini Flash-tier standard pricing per 1M tokens, read on Google's pricing page August 24, 2026
ModelInput nowOutput nowInput from Jan 1Output from Jan 1Cached input
Gemini 3.7 Flash$0.75$3.75$1.50$7.50$0.075
Gemini 3.6 Flash$0.75$3.75$1.50$7.50$0.075
Gemini 3.5 Flash$1.50$9.00$1.50$9.00$0.15
Gemini 3.5 Flash-Lite$0.30$2.50$0.30$2.50n/a

Two things fall out of that table immediately. Gemini 3.5 Flash is now the most expensive Flash model Google sells and the oldest of the three, which is a combination nobody should be paying for. And 3.6 Flash, which benchr recorded at $1.50 / $7.50 on July 22, has quietly been cut in half to match the newcomer.

The January 1 math, on a real workload

Take a moderate agent service: 40M input tokens and 12M output tokens a month, no caching. On 3.7 Flash today that's $30 + $45 = $75 a month. On January 1, with identical traffic and identical code, it becomes $60 + $90 = $150 a month. Nothing about your system changed. The sticker did.

Caching moves that line but doesn't remove it. Cached input is $0.075 per million now and $0.15 from January. Cache storage runs $0.50 per million token-hours now and $1.00 later, so the storage side doubles too. That matters if you keep large prefixes warm. Run the numbers on your own volumes in the API cost calculator rather than trusting a round figure, and run them twice: once at today's rate, once at January's.

One month of a 40M in / 12M out agent, at each published rate

Monthly token cost calculated from Google's published per-token rates, read August 24, 2026. Standard tier, uncached, no grounding or cache-storage charges. Bar lengths are proportional to the dollar figures shown.

Gemini 3.5 Flash-Lite, no expiry date
$42
Gemini 3.7 Flash, today
$75
Gemini 3.7 Flash, from January 1, 2027
$150
Gemini 3.5 Flash, unchanged either way
$168

What Google says changed, and what it didn't publish

The release note credits 3.7 Flash with substantial improvements across software engineering, web development, and agentic workflows. The model page fills in the surface: text, image, video, audio, and PDF input with text output, a 1,048,576-token input limit and 65,536-token output limit, caching, code execution, file search, function calling, Search and Maps grounding, structured outputs, URL context, and Computer Use in preview. Thinking runs at low, medium, or high — there's no minimal level on this one. Image generation, audio generation, and the Live API aren't supported.

What's missing is any benchmark table. Google published none with the release note and none on the model page, which is the same gap that followed the 3.6 Flash launch. That's worth stating plainly rather than papering over: the claim that 3.7 is stronger than 3.6 currently rests entirely on Google's own description of it. benchr's capability profiles on the models page are editorial judgments, not provider results, and they shouldn't be read as filling that gap either.

So the decision splits cleanly. On price, the two models are indistinguishable and the choice is free. On capability, you have a vendor claim and no published numbers, which means the tiebreaker has to come from your own replay of real traffic.

What to do this week

  1. If you're on Gemini 3.5 Flash, move. You're paying $1.50 / $9.00 for the oldest model in the family while two newer ones sit at $0.75 / $3.75. Pin gemini-3.7-flash, replay a production-shaped trace, then compare accepted-result cost rather than raw token counts.
  2. If you're on 3.6 Flash, there's no rush and no penalty. The price is identical, so migrate when your test window allows. Read the 3.6 Flash pricing page for the tier-by-tier breakdown if you're staying put for now.
  3. Put January 1, 2027 in the budget document, not just the calendar. Model the doubled rate against next year's projected volume. If the doubled number breaks the unit economics, the time to find that out is while you still have four months to re-architect, not in the first week of January.
  4. Check whether Flash-Lite covers the work. At $0.30 / $2.50 with no expiry attached, Gemini 3.5 Flash-Lite is cheaper than the promotional Flash rate on output and stays cheaper afterward. For classification, extraction, and subagent fan-out it's often enough.

How this compares outside Google

August was a discount month across the board, so the Flash rate doesn't sit still against rivals either. OpenAI cut GPT-5.6 Sol to $4 / $20 on August 21, also on promotional terms, this time with a November 21 floor rather than a December 31 one. GPT-5.6 Luna remains the cheapest capable option on that side at $0.20 / $1.20. Against Luna, Gemini 3.7 Flash costs more on both sides of the meter; what it offers instead is a million-token window, native multimodal input, and Google's grounding stack. Those are architecture reasons to choose it rather than price reasons, and they'll still be true in January when the price reason expires.

FAQ

How much does Gemini 3.7 Flash cost?

Google lists $0.75 per million input tokens and $3.75 per million output tokens on the standard paid tier through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. Context caching is $0.075 per million now and $0.15 later, with storage at $0.50 then $1.00 per million token-hours. Batch and Flex are half of standard, Priority is $1.35 / $6.75, and a free tier is available.

Is Gemini 3.7 Flash better than 3.6 Flash?

Google describes it as substantially improved for software engineering, web development, and agentic workflows, but publishes no benchmark table for either model. The price is identical, so there's no cost penalty in choosing 3.7 — the capability difference has to be confirmed on your own workload.

What are Gemini 3.7 Flash's token limits?

The official model page lists a 1,048,576-token input limit and a 65,536-token output limit for gemini-3.7-flash. It accepts text, image, video, audio, and PDF input and returns text only.

Will the introductory price be extended?

Google's pricing page states the higher rate starts January 1, 2027 and does not describe any extension. Plan for the published date. If it changes, the change will appear on the pricing page and in benchr's price history.

Sources

  1. Google AI for Developers, Gemini API release notes — establishes the August 13, 2026 general-availability date and Google's description of what 3.7 Flash improves. Verified August 24, 2026.
  2. Google AI for Developers, Gemini API pricing — source for every rate in this article, including the December 31, 2026 promotional end date and the January 1, 2027 rates for both 3.7 and 3.6 Flash. Verified August 24, 2026.
  3. Google AI for Developers, Gemini 3.7 Flash model page — source for the token limits, supported input and output types, capability list, and the absence of a benchmark table. Verified August 24, 2026.
  4. OpenAI, OpenAI API pricing — source for the GPT-5.6 Sol and Luna rates used as the cross-provider anchor. Verified August 24, 2026.

Changelog

August 24, 2026: Published after a re-read of Google's pricing table showed Gemini 3.6 Flash cut to match the Gemini 3.7 Flash introductory rate, with both reverting January 1, 2027.