The playbook
An idea to a narrated product clip
Stills first, motion second, voice last - the order that keeps control where you still have it.
3 stages A day, most of it iterating on stills Gemini API · OpenAI API Complete
Every stage names what goes in, what to do, why it sits here and what comes out, and the chain says where it breaks.
benchr has not run these capabilities itself. Each stage points at a record describing what a provider documents, read on the date shown.
- You start with
- A photograph of the real product, the brand assets, and the exact words that have to appear on screen
- Why this stage
- Every clip is animated from one of these stills, so the model choice belongs here: the record's own pre-flight is to check the resolution ceiling before you pick the model, and the Flash Lite image model supports 1K only. Wording is settled here too - legible text is described as a strength, not a guarantee, and long strings and small type stay the weak case.
- Do this
- Generate the key frames with the exact copy quoted in the prompt and brand references attached.
- You end with
- Approved stills with correct wording
What goes wrong here
- You start with
- One approved still per shot
- Why this stage
- The still goes in as the starting frame, which the record calls the reliable way to control composition; the same record calls this the least steerable capability on the list, so budget for iteration. The model choice lands here too: the documentation names Gemini Omni Flash as the default recommendation, and native audio is documented for Veo 3.1 only, not for the default model - so whether the clips carry sound of their own is decided before there is a narration track.
- Do this
- Animate each approved still as a separate short clip - one camera move, one action each - and generate several takes of each. Expect to discard most; selection is the work.
- You end with
- Short clips you can cut together, chosen from several takes
- You start with
- The lines to be spoken, written to time
- Why this stage
- A live session earns its place when the read has to be heard and corrected while it is happening: audio goes in, audio comes back, and the turn-taking is handled for you. Design the interruption behavior before the prompt - the record calls barge-in the difference between a demo and something people will use. The price is a session protocol, not a request-response endpoint: reconnection, state and audio buffering are yours, and the latency and quality trade-off is explicit in the documentation and cannot be optimized away, only chosen. For a read that needs no live performance, capabilities.json has no text-to-speech record to send you to.
- Do this
- Use the realtime voice stack for narration only if you need a live take; otherwise script it and record once.
- You end with
- A narration track