Gemini 3.5 Flash-Lite: the $0.30 subagent target has a migration cost
Flash-Lite's $0.30 input rate suits large queues. Before switching, check the thinking defaults and request fields Google changed.
By the benchr team · · Changelog · Provider-published facts rechecked against the official sources on July 28, 2026
GOOGLE · GEMINI3.5FLASH-LITE · VOLUME
Benchr model field plateGemini 3.5 Flash-LiteHigh throughput · migration rules
Editorial imageA BenchR editorial illustration of high-volume routing: low cost changes the queue, while migration rules still matter.
Input / 1M$0.30Output: $2.50
Context1M64K max output
StatusGAProduction-ready
Released21 Jul2026
Gemini 3.5 Flash-Lite became generally available on July 21. The appeal is obvious: it is the cheapest member of Google's 3.5 family. The harder question is which calls are simple enough to send there.
Give it queues, not every job
Google highlights high-volume parsing, document extraction, structured JSON, and autonomous subagents. Those are good first targets. For an agent that needs deeper planning, Google suggests increasing the thinking level. Do not assume the cheapest default will preserve the same plan quality.
The $0.30 rate needs the whole request
Flash-Lite costs $0.30 input and $2.50 output per million tokens. Output and thinking tokens are billed as output, tools change the request pattern, and a 1M-token window can still encourage expensive history. Price a completed job, not the input line by itself.
Old request controls can break the move
Google says newer 3.x Flash releases deprecate sampling parameters such as temperature, top_p, and top_k, and no longer support prefilled model turns. Replay a sample of production requests before changing the model ID. A cheaper subagent that returns 400s is not a cheaper system.