The question enterprise leaders have been asking for two years finally has a data-driven answer: OpenAI just published proof that its own AI agent resolves 75% of inbound support issues without a human — and cut human handoffs by 15 percentage points in just 10 days.
That's not a benchmark. That's not a proof-of-concept. That's a production number from OpenAI's own phone support line, running on the same platform it's now selling to enterprise customers as OpenAI Presence.
If you're a CIO, CTO, or CFO evaluating whether enterprise AI agents are ready for mission-critical work, this announcement changes your calculus. Here's what you need to know.
What OpenAI Presence Actually Is
OpenAI Presence isn't another chatbot platform. It's a battle-tested enterprise deployment framework for AI agents that combines model reasoning with the governance layer most enterprise AI projects skip entirely — and that's exactly why most fail.
The platform launched July 22, 2026, in limited general availability. It's available to enterprise customers through a deployment led by OpenAI's Forward Deployed Engineers (FDEs) and select global systems integrators. It is not yet self-serve.
That last point is worth sitting with: OpenAI is deliberately not making this a click-to-deploy product. They've built a white-glove deployment model where their engineers partner with your team to identify high-value workflows, connect the necessary knowledge and systems, establish permissions and policies, test the agent, and bring it into production. That approach is either a feature or a limitation, depending on your situation.
Each Presence deployment starts with a specific job — resolving billing issues, supporting insurance claims, handling employee IT service requests. The agent gets only the knowledge and system access required for that job. Your company sets the policies: what the agent can do, when it needs approval, and when a human should take over.
After launch, production sessions and escalations surface gaps. OpenAI's Codex integration then proposes updates that your team tests and approves before rollout. That's the improvement loop that dropped human handoffs 15 percentage points in 10 days.
Why Most Enterprise AI Agents Fail
Before evaluating Presence specifically, it's worth understanding the broader failure rate this product is designed to address.
Gartner projects that over 40% of enterprise AI agent projects will be canceled by 2027. Only 41% of AI agent deployments achieve positive ROI within the first year. And 19% never reach payback at all — often due to poor evaluation, governance gaps, and unmeasured rework.
Those numbers should give any enterprise leader pause. The technology clearly works — OpenAI's own data proves it. The failure pattern isn't model capability. It's deployment discipline.
In conversations with enterprise architecture leaders across industries, I hear the same root causes repeatedly: agents that hallucinate in production, no clear escalation policy when the agent hits an edge case, no mechanism to update agent behavior as products and policies change, and no governance framework that passes a legal or compliance review. Organizations underestimate the total cost of AI agent deployment by 40–60%, with implementation and tuning costs running 3–5x the software license cost.
OpenAI Presence is explicitly designed to address this failure mode. The governance components — policies and standard operating procedures, guardrails, approved actions, simulations, evaluation tools — are the parts that enterprise teams consistently underinvest in when building their own agent stacks.
For CTOs: What's Under the Hood
From a technical architecture standpoint, Presence has several components worth understanding.
Policies and standard operating procedures define what the agent can and cannot do, codified in a way that the model reasons against. This isn't prompt engineering — it's a structured policy layer that sits between user requests and agent actions.
Guardrails intervene when an interaction moves outside your company's defined boundaries. This is the control plane that separates Presence from a raw API deployment.
Approved actions define what the agent can execute autonomously versus what requires human sign-off. This is the governance primitive that makes the difference between an agent that's useful in production and one that creates liability.
Simulations and graders test the agent against common requests, edge cases, and higher-risk scenarios before it reaches users. These evaluate whether the agent reached the right outcome, followed policy, used tools correctly, and escalated appropriately.
Codex-powered improvement loop analyzes production sessions and escalations to identify where the agent is underperforming, then proposes specific updates. Teams can test each proposed change against the production version before approving a controlled rollout.
Today, Presence supports voice and chat channels. Customer use cases at launch include customer support, outbound sales, and high-risk internal workflows. The agent handles the full interaction stack: understanding the request, verifying the customer, looking up account information, applying company policy, and taking an approved action.
BBVA is exploring AI-powered voice support for everyday banking needs in Mexico. SoftBank is testing natural Japanese-language customer conversations. IAG is exploring timely support during high-demand events like severe weather. These aren't technology companies building edge cases — they're large, regulated enterprises running complex customer interactions.
For CFOs: The Cost Math
The ROI case for enterprise AI agents is becoming clearer as production data accumulates, and the numbers are compelling.
AI agents can handle a contained support ticket for approximately $0.46, compared to $4.18 for a human agent — a 9x cost reduction per interaction. The cost per interaction for an AI agent ranges from $0.25 to $0.50, an 85–90% reduction from the $3.00 to $6.00 cost of a human interaction.
Enterprises deploying agentic AI report an average ROI of 171%, with US-based enterprises averaging 192%. The median payback period for AI agent deployments is 5.1 months overall, and 4.1 months specifically for customer service agents.
But here's the nuance that separates the 41% that achieve positive ROI in year one from the 59% that don't: vendor-deployed agents achieve positive ROI 2.4x faster than custom-built solutions.
That's the strategic argument for Presence over a build-your-own approach. Your engineering team can absolutely build agent infrastructure — but the deployment expertise, the governance tooling, the improvement loop, and the battle-tested failure modes take years to develop. OpenAI is selling that institutional knowledge, not just a model.
At 75% autonomous resolution, if your customer support volume is 10,000 tickets per month at $4.18 per human interaction, a Presence deployment potentially shifts 7,500 of those tickets to $0.46 each. That's a monthly cost reduction from $41,800 to about $30,750 — saving roughly $11,000 per month, or $132,000 per year, just on ticket cost. The math scales quickly at enterprise volumes.
The Competitive Landscape Shift
OpenAI Presence enters a market that Salesforce Agentforce has been building for 18 months. Agentforce reached $800 million in annual recurring revenue, up 169% year-over-year — proof that enterprise demand for AI agents is real and accelerating.
OpenAI's play is positioned differently. Salesforce Agentforce is a platform that integrates with Salesforce CRM data and workflows. Presence is a deployment methodology plus AI infrastructure that can connect to any company system. If your enterprise isn't deeply embedded in Salesforce, Presence potentially offers more flexibility.
ServiceNow has its own agentic AI offering for IT service management. Zendesk has built AI agent capabilities into its customer service platform. What OpenAI is betting is that its frontier model capability, combined with a white-glove deployment model, can outcompete purpose-built vertical tools by delivering better agent reasoning on complex, open-ended requests.
The decision framework for enterprise buyers isn't simply "which AI agent platform?" It's "do we want a platform that integrates deeply into our existing enterprise software stack, or do we want a deployment model that starts from OpenAI's frontier model capability and connects to our systems?"
That's a meaningfully different procurement conversation than it was 12 months ago.
What Enterprise Leaders Should Do Now
This isn't a "wait and see" situation. Fifty-one percent of enterprises already have AI agents in production. The competitive differentiation window is narrowing.
Here's the practical roadmap.
Identify your highest-volume, highest-cost repetitive workflows. Customer support, IT help desk, HR benefit queries, and billing inquiries are the obvious candidates. Calculate current cost per interaction and resolution rate. That's your ROI baseline.
Evaluate your governance readiness before evaluating vendors. The enterprises that fail with AI agents aren't failing because the technology doesn't work — they're failing because they don't have clear policy frameworks, escalation protocols, and evaluation methodologies. Presence provides these, but your team still needs to define the policies. That work takes time.
Run the build vs. buy analysis honestly. If your engineering team has 6–12 months to build governance tooling, evaluation frameworks, and improvement loops, a custom stack may make sense. If you need production results in 3–6 months, the 2.4x faster ROI of vendor-deployed solutions changes the math significantly.
Ask OpenAI hard questions about the FDE model. Presence is not self-serve. Your deployment will be led by OpenAI engineers and select systems integrators. How does that work at your scale? What happens when you want to modify policies 18 months into deployment? What's the ongoing support model? These are reasonable questions for any enterprise that takes IT vendor management seriously.
Pilot on a contained use case. The OpenAI announcement specifically mentions starting with "a specific job — such as resolving billing issues, supporting insurance claims, or resolving employee IT service requests." That's not a coincidence. It's deployment wisdom. Start with a bounded, measurable workflow where success criteria are clear.
The Bottom Line
OpenAI Presence is the most credible enterprise AI agent product I've seen announced — because OpenAI is eating its own cooking. They built it, ran it in production on their own support line, measured it, and published the results: 75% autonomous resolution, 15 percentage points fewer human handoffs in 10 days.
Those numbers aren't theoretical. They're operational. And they're the kind of data that moves enterprise procurement conversations from "interesting pilot" to "this deserves a business case."
The model is just one ingredient. The governance layer, the improvement loop, and the deployment discipline are what separate enterprise-grade AI agents from demos that never make it to production. OpenAI is betting that its combination of frontier models and deployment expertise can unlock that value for enterprises that have struggled to get there on their own.
Whether Presence is right for your organization depends on your existing vendor relationships, your engineering capacity, your timeline, and your risk tolerance. But every enterprise leader evaluating customer experience, IT operations, or internal workflow automation should have this conversation with their teams this quarter.
The 40% of projects that get canceled don't fail because AI agents don't work. They fail because the governance and deployment discipline wasn't there. OpenAI is selling that discipline, not just the model.
For more on enterprise AI agent deployments, governance frameworks, and vendor comparisons, explore:
- The Real Cost of Enterprise AI Agents — ROI frameworks and total cost of ownership
- Salesforce Agentforce vs. Custom AI: The CIO's Framework — Platform comparison
- Why 40% of Enterprise AI Projects Fail (And How to Not Be That Statistic) — Governance playbook
