Presence answers a different buying question from the one covered by an API pricing guide. The API gives a team model access and primitives. Presence sells a larger operating system around an agent: policies, standard operating procedures, guardrails, approved actions, simulations, evaluations, change control, and specialists who help move a chosen workflow into production.
What OpenAI is actually selling
The launch page starts with a job, not a general assistant. OpenAI's examples include billing resolution, insurance claims, and employee IT requests. The agent receives only the knowledge and system access needed for that job. The customer decides what it may do, which actions need approval, and when a person takes over.
| Layer | Presence includes | Buyer should verify |
|---|---|---|
| Conversation | Real-time voice and chat experiences | Languages, latency, telephony, regions, accessibility |
| Knowledge and tools | Connections to company information and systems | Connector scope, identity, least privilege, audit logs |
| Control | Policies, SOPs, guardrails, approvals, escalation rules | Who authors rules, versioning, override and rollback |
| Quality | Simulations, graders, evaluation tools, production signals | Held-out sets, false-pass rate, incident review, exports |
| Improvement | Codex-assisted investigation and proposed updates | Approval boundary, regression testing, change history |
| Delivery | OpenAI FDEs and selected systems integrators | Ownership, support levels, exit plan, custom-work cost |
This packaging matters because most failed agent projects do not fail at the model call. They fail at permission design, incomplete edge cases, weak evaluation, brittle integrations, or an unsafe escalation path. Presence puts those operational pieces inside one commercial engagement. That can reduce assembly work, but it also makes the service harder to compare with a token price.
How to read the 75% and 15-point claims
OpenAI says Presence powers its English-language phone support line, resolves 75% of inbound issues without human assistance, and reduced human handoffs by 15 percentage points in ten days through its improvement loop. Those numbers are useful proof that the product is operating in a live environment. They are not an independent benchmark.
The announcement does not publish the case mix, evaluation denominator, exclusion rules, customer-satisfaction result, repeat-contact rate, escalation severity, or cost per resolved issue. A buyer should not transfer the 75% figure to insurance, banking, or IT support without running a representative pilot. The right comparison is not “agent versus human” in the abstract. It is current workflow versus Presence on the same requests, policies, languages, and service-level targets.
Presence versus building on the API
| Decision | Presence is the stronger candidate when… | API-first is the stronger candidate when… |
|---|---|---|
| Time to controlled production | You need deployment expertise and an operating process | You already own agent infrastructure and evaluations |
| Customization | The workflow fits the supported managed product | You need unusual orchestration or product-level control |
| Procurement | Enterprise contracting and services are acceptable | You need public usage pricing and self-serve iteration |
| Operating ownership | You want OpenAI or an integrator in the change loop | Your team must own every release and dependency |
| Portability | Managed outcomes matter more than provider portability | A model- or provider-switching layer is a requirement |
Presence does not replace the API. OpenAI explicitly says voice customers will continue to have access to frontier models through the API. The two offers sit at different layers. Teams comparing foundation models can still use the voice-model comparison; Presence is a build-versus-buy decision about the system around those models.
The commercial gaps matter
The public launch page does not disclose price, minimum commitment, implementation duration, supported countries, language coverage beyond the cited English and Japanese examples, detailed service levels, or a complete list of connectors. It also does not say which model versions power each deployment. These are not footnotes. They determine whether Presence is a product, a services engagement, or a mixture of both for a particular buyer.
Privacy needs the same separation. OpenAI's business-data page says business inputs and outputs are not used for training by default, and its API data-control documentation describes retention, regional processing, and zero-data-retention options for eligible API customers. Presence is a distinct managed product. A buyer should obtain the exact contractual retention, residency, subprocessors, access, logging, and deletion terms that apply to the Presence deployment rather than assuming every API control carries over unchanged.
Put the evidence in the contract, not the demo
OpenAI's separate data-usage guidance repeats that business and API inputs and outputs are excluded from training by default unless an organization opts in. That is useful background, but Presence is not named in that product list. The order form and data-processing terms should identify the service explicitly. Ask for a written system boundary: every model, telephony service, connector, log store, support-access path, and subprocessor that can receive customer content. The same annex should state retention by data class, deletion timing, permitted support access, incident notification, export format, region, and the controls that survive termination.
Commercial evidence deserves the same treatment. Require one sheet that separates implementation fees, recurring platform charges, model usage, telephone minutes, integration work, support, and change requests. Tie service levels to observable events—availability, response latency, failed tool calls, escalation delivery, and recovery—not to a vague promise that the agent will “improve.” A polished demonstration proves a path works once; the contract must describe who owns failures and changes after launch.
A procurement test that produces an answer
- Choose one bounded workflow. Freeze the request types, actions, policies, languages, and escalation conditions.
- Build a replayable evaluation set. Include ordinary cases, policy conflicts, identity failures, tool timeouts, malicious instructions, and requests that must reach a person.
- Define release gates before the pilot. Set thresholds for correct resolution, unsafe action, false approval, escalation quality, latency, repeat contact, and cost per completed case.
- Compare against two baselines. Use the existing human-assisted process and the narrowest credible API-built alternative. Include implementation and ongoing review labor.
- Test change control. Modify a policy, break a connector, and introduce a new edge case. Measure detection, proposed fix quality, regression results, approval, rollout, and rollback.
- Write the exit plan. Identify which prompts, policies, evaluation sets, logs, integration code, and data exports remain usable if the service is replaced.
Benchr's judgment
Presence is most interesting because OpenAI is no longer presenting the model as the product. The offer is the controlled loop around the model: permissions, evaluation, escalation, and continuous changes. That is closer to how enterprises actually buy reliable automation.
It is too early to call the service a value winner. Without public pricing, service-level detail, and comparable outcome data, the responsible conclusion is narrower: Presence deserves a pilot when a company has a high-volume voice or chat workflow, clear policies, measurable outcomes, and enough risk that deployment support matters. Teams seeking a transparent cost curve, self-serve access, or provider portability should keep an API-first baseline in the evaluation.