Gemini 3.5 Flash-Lite: the $0.30 subagent target has a migration cost

Flash-Lite's $0.30 input rate suits large queues. Before switching, check the thinking defaults and request fields Google changed.

By the benchr team · · Changelog · Provider-published facts rechecked against the official sources on July 28, 2026

Many small coloured task tiles converge into one narrow fast processing lane.
Benchr model field plate Gemini 3.5 Flash-Lite High throughput · migration rules
Editorial imageA BenchR editorial illustration of high-volume routing: low cost changes the queue, while migration rules still matter.
Input / 1M$0.30Output: $2.50
Context1M64K max output
StatusGAProduction-ready
Released21 Jul2026

Gemini 3.5 Flash-Lite became generally available on July 21. The appeal is obvious: it is the cheapest member of Google's 3.5 family. The harder question is which calls are simple enough to send there.

Give it queues, not every job

Google highlights high-volume parsing, document extraction, structured JSON, and autonomous subagents. Those are good first targets. For an agent that needs deeper planning, Google suggests increasing the thinking level. Do not assume the cheapest default will preserve the same plan quality.

The $0.30 rate needs the whole request

Flash-Lite costs $0.30 input and $2.50 output per million tokens. Output and thinking tokens are billed as output, tools change the request pattern, and a 1M-token window can still encourage expensive history. Price a completed job, not the input line by itself.

Old request controls can break the move

Google says newer 3.x Flash releases deprecate sampling parameters such as temperature, top_p, and top_k, and no longer support prefilled model turns. Replay a sample of production requests before changing the model ID. A cheaper subagent that returns 400s is not a cheaper system.

Provider-published facts; documented gaps remain gaps
FieldVerified record
API model IDgemini-3.5-flash-lite
Input / output price$0.30 / $2.50 per 1M
Context / output1M / 64K tokens
PositioningHigh-throughput subagents and document extraction

Frequently asked

What is Gemini 3.5 Flash-Lite for?

Google positions it for high-throughput execution, document extraction, structured data work, and subagents.

How much does Gemini 3.5 Flash-Lite cost?

Google lists $0.30 input and $2.50 output per million tokens on its paid standard tier.

What must be checked before migrating?

Google documents deprecated sampling parameters and disallows prefilled model turns for the newer release family.

Changelog

  • July 28, 2026 — Published after reviewing the official provider sources and recording unreported fields as gaps.

References

  1. Official release announcement or release notes: https://ai.google.dev/gemini-api/docs/changelog
  2. Official pricing documentation: https://ai.google.dev/gemini-api/docs/latest-model
  3. Official model documentation: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite