GPT-Realtime-2.1 mini: the lower-cost voice model still has an audio bill

Mini keeps the 128K window and cuts every listed rate. Audio still dominates the calls you need to measure.

By the benchr team · · Changelog · Provider-published facts rechecked against the official sources on July 28, 2026

Small audio modules passing through a narrow efficient processing lane.
Benchr model field plate GPT-Realtime-2.1 mini Compact voice workload
Editorial imageA BenchR editorial illustration of a smaller realtime workload designed around a compact cost shape.
Text / 1M$0.60Output: $2.40
Audio / 1M$10Output: $20
Context128K32K max output
Released6 Jul2026

OpenAI announced GPT-Realtime-2.1 mini on July 6 as the cheaper companion to GPT-Realtime-2.1. The name is accurate, but it can tempt you into comparing one price ratio and calling the budget done.

Mini is cheaper in every listed modality

Text costs $0.60 input and $2.40 output per million tokens. Audio is $10 input and $20 output, while image input is $0.80. Those rates sit well below the larger model. Your bill still depends on how much of each modality a completed call uses, not the text line alone.

The big window still needs a history rule

Mini lists the same 128K context and 32K maximum output as GPT-Realtime-2.1. That can preserve a long call state, but retaining every turn can raise cost and bury useful details. Decide what gets summarized, then replay calls with the same tools, speaker changes, and recovery steps you expect in production.

Make the larger model defend its premium

OpenAI calls mini a distilled reasoning model for faster, lower-cost voice interactions. Pair it with the larger model on the same calls. Compare completed tasks, interruption repairs, tool accuracy, and cost per successful conversation. Promote only the cases where the larger model wins enough to matter.

Provider-published facts; documented gaps remain gaps
FieldVerified record
API model IDgpt-realtime-2.1-mini
Text / audio price$0.60 / $2.40; $10 / $20 per 1M
Image input$0.80 per 1M tokens
Context / output128K / 32K tokens

Frequently asked

How is GPT-Realtime-2.1 mini priced?

OpenAI lists $0.60/$2.40 text input/output, $10/$20 audio input/output, and $0.80 image input per million tokens, with separate cached-input fields where shown.

Does mini have the same context as GPT-Realtime-2.1?

Yes. The official pages list 128K context and 32K maximum output for both.

Is mini a text-only model?

No. OpenAI lists text, audio, and image billing categories for the realtime model.

Changelog

  • July 28, 2026 — Published after reviewing the official provider sources and recording unreported fields as gaps.

References

  1. Official API release record: OpenAI API changelog
  2. Official model and pricing documentation: GPT-Realtime-2.1 mini model page