Gemini 3.1 Flash TTS Preview: steerable speech with preview boundaries
Google publishes the input and output limits and supports Batch. The model card still leaves the price blank.
By the benchr team · · Changelog · Provider-published facts rechecked against the official sources on July 28, 2026
GOOGLE · GEMINI3.1FLASH · TTS · PREVIEW
Benchr model field plateGemini Flash TTSText in · directed sound out
Editorial imageA BenchR editorial illustration of the transformation from a written signal to directed audio output.
Input limit8,192Text tokens
Output limit16,384Audio tokens
Batch APIYesDocumented
Launched15 Apr2026
Gemini 3.1 Flash TTS Preview launched April 15. It is a focused speech generator, not a general Gemini model that happens to speak.
Build around synthesis
The card lists text input, audio output, audio generation, and Batch support. It does not list function calling, code execution, Live API, grounding, structured output, or thinking. Keep the workflow simple: prepare text, generate audio, review the result.
Long scripts need sections
The 8,192-token input and 16,384-token output limits matter for narration and multi-speaker work. Split a long script on meaningful boundaries, keep the content map outside the model, and listen for awkward joins. One giant prompt is not a production plan.
The preview has no published rate
Google labels the model a preview, and the checked card lists no per-token price. That does not mean free. Keep cost as an unknown, watch the model and pricing pages, and get a current quote before committing a production budget.