GPT-Realtime-2.1: evaluate the conversation loop, not a text-only price card

This voice model is built around interruptions, tools, and recovery. Its audio tokens also make the text price a poor budget shortcut.

By the benchr team · · Changelog · Provider-published facts rechecked against the official sources on July 28, 2026

A continuous audio loop reconnecting around a central tool-routing point.
Benchr model field plate GPT-Realtime-2.1 Voice loop · tools · interruption
Editorial imageA BenchR editorial illustration of the conversation loop: interruption, recovery, and a tool call belong in one test.
Text / 1M$4Output: $24
Audio / 1M$32Output: $64
Context128K32K max output
Released6 Jul2026

GPT-Realtime-2.1, announced July 6, is not a text model with a microphone attached. The interesting part is the loop around speech: what happens when a caller interrupts, spells an account number, triggers a tool, or talks over background noise.

Test the messy call, not the studio demo

OpenAI says the update improves alphanumeric recognition, silence and noise handling, interruptions, instruction following, and tool use. Build those failures into the test. Include barge-in, a noisy room, spelling-sensitive names, a slow tool call, and a recovery turn. A clean scripted transcript skips the work this model is meant to do.

The audio line is the budget

Text costs $4 input and $24 output per million tokens. Audio costs $32 input and $64 output. A text-only spreadsheet can make this model look cheap while a call replay tells a different story. Track text, audio, and cached input separately, then price a completed call rather than a token category.

Do not keep every turn forever

A 128K context window and 32K output ceiling give you room, not a call-history policy. Decide what must survive an interruption, when tool results go stale, and when the transcript should be summarized. Score an end-to-end trace for turn-taking and task completion.

Provider-published facts; documented gaps remain gaps
FieldVerified record
API model IDgpt-realtime-2.1
Text pricing$4 input / $24 output per 1M
Audio pricing$32 input / $64 output per 1M
CapabilitiesSpeech-to-speech, configurable reasoning effort, tool use

Frequently asked

What is GPT-Realtime-2.1 for?

OpenAI describes it as a speech-to-speech reasoning model with tool use for complex voice-agent workflows.

What context window does GPT-Realtime-2.1 have?

The official model page lists a 128,000-token context window and a 32,000-token maximum output.

Why are its prices hard to compare with a text model?

OpenAI bills text and audio tokens in separate categories, so a text-only comparison omits audio input and output.

Changelog

  • July 28, 2026 — Published after reviewing the official provider sources and recording unreported fields as gaps.

References

  1. Official API release record: OpenAI API changelog
  2. Official model and pricing documentation: GPT-Realtime-2.1 model page