OpenAI introduced GPT-Live-1 on July 8 for ChatGPT Voice and said API access was planned. That last word matters. The product exists, but the developer contract does not.
You can use it in ChatGPT, not your stack
GPT-Live-1 is the full-duplex Voice model rolling out to eligible ChatGPT Go, Plus, and Pro users. That tells you where the product can be tried. It does not give you a stable model string, an API endpoint, or a production rate card.
Full duplex is the feature
OpenAI says the Live models can listen and speak in the same interaction, then hand deeper work to other OpenAI models. The announcement does not publish a context window, output cap, benchmark table, or API price. A ChatGPT plan price cannot fill those blanks.
Run the product trial now; plan the integration later
A support or research team can test whether the Voice experience helps a user today. A developer team still cannot estimate API spend or promise endpoint behavior from this launch post. Write down the endpoint, billing, privacy, tool, and fallback requirements, then wait for documentation that answers them.
| Field | Verified record |
|---|---|
| Audience | ChatGPT Voice |
| Interaction | Full duplex |
| Public API ID | Not published in checked announcement |
| Public API price / limits | Not published in checked announcement |
Evaluate the conversation, not just the voice
OpenAI describes continuous interaction: the system can keep listening while it speaks and decide whether to pause, interrupt, or invoke a tool. Test it with timed conversations, not isolated audio samples. Measure false interruptions, recovery after overlap, response to silence, background-noise handling, and whether a task remains understandable after a delegated answer returns.
| Scenario | Evidence to capture | Do not infer |
|---|---|---|
| User interrupts | Stop time and context recovery | API latency guarantee |
| Long pause | Whether the system waits or cuts in | Turn-detection setting |
| Search or reasoning | Handoff clarity and returned result | One-model capability |
| Non-English session | Accent, fluency, misunderstanding | Universal language parity |
GPT-Live is a system boundary
The launch says deeper work can be delegated to a frontier model while GPT-Live maintains the conversation. That means the observed answer is a system result, not a clean benchmark of the voice model alone. For procurement, separate continuous interaction, delegated intelligence, search or tool use, and the ChatGPT interface in the evaluation notes.
Known launch limits belong in the test plan
OpenAI says some languages may have a non-native accent or fluency gaps, and that voice with video or screen sharing is not supported at launch. Those are concrete boundaries. A team considering multilingual support or visual assistance should test the current product without assuming the announced future additions are already available.
API readiness checklist
Do not freeze an architecture until OpenAI publishes the model identifier, transport, billing units, rate limits, tool contract, data controls, and failure behavior. The launch post says API access is planned soon; “planned” is a watch signal, not an interface contract.
Run a moderated session with a written rubric
Use the same opening instruction and task sequence for every participant. Include a correction, an interruption, a period of silence, a noisy moment, and one request that triggers deeper work. The moderator should timestamp when the user starts and stops speaking, when the system yields, and whether the participant needs to repeat information. Afterward, ask what the user believed was happening during any handoff.
Separate conversation flow from answer quality. A fast, natural exchange may still return a weak answer, while a strong delegated answer may arrive after a confusing pause. Record both. Also distinguish features of ChatGPT—memory, visual cards, search, file support—from the behavior attributed specifically to GPT-Live. The launch describes an integrated product, so the test report should not pretend every observed capability belongs to one model.
Before exposing voice to customers, define consent, transcript retention, escalation, and a human-transfer path. Test names, account numbers, addresses, and other details that are costly to mishear, without using real customer data in the pilot. These safeguards are product requirements today; they should not wait for a future API price sheet.