OpenAI Presence: 75% Issue Resolution, No Humans Needed

OpenAI Presence deploys production-ready AI agents that resolve 75% of enterprise issues without human handoffs. What CIOs and CFOs need to know now.

By Rajesh Beri·July 24, 2026·11 min read
Share:
THE DAILY BRIEF
OpenAIEnterprise AIAI AgentsCustomer Service AIAgentic AI
OpenAI Presence: 75% Issue Resolution, No Humans Needed

OpenAI Presence deploys production-ready AI agents that resolve 75% of enterprise issues without human handoffs. What CIOs and CFOs need to know now.

By Rajesh Beri·July 24, 2026·11 min read

When OpenAI announced Presence last week, the headline number was 75% — the share of inbound support calls resolved without any human assistance. That stat is real. But the bigger story is what it signals about where enterprise AI is heading, and why the window for building your own solution is getting shorter.

Most enterprise AI deployments I've seen over the past two years share a common problem: the first pilot works great. The second one requires twice the infrastructure. By the fourth or fifth deployment, you're drowning in custom prompt chains, brittle integrations, and a team of engineers who spend more time firefighting than building. The governance problem — making agents behave consistently, safely, and adaptably — has been the hard part nobody talks about at conferences.

OpenAI Presence is a direct answer to that problem. And based on their early production results, it's working.

What OpenAI Presence Actually Is

Presence is OpenAI's enterprise product for deploying governed AI agents across voice and chat workflows. It launched July 22, 2026, and is available immediately through a limited general availability program — but it's not self-serve. You need OpenAI Forward Deployed Engineers (FDEs) and select global systems integrators to set it up.

The core idea: instead of asking your engineering team to stitch together models, APIs, security controls, and evaluation tooling, OpenAI packages that entire infrastructure and deploys it alongside you.

Each Presence deployment starts with a specific job — resolve billing issues, handle insurance claims, answer IT service requests. The agent receives only the knowledge and system access needed for that job. Your organization sets the policies: what the agent can do independently, which actions need human approval, and when to escalate entirely.

The product includes what OpenAI calls the complete operational stack: policies and standard operating procedures, guardrails, approved actions, simulation testing, evaluation tools, and a Codex-powered improvement process that monitors production behavior and proposes updates. Teams review proposed changes, test them against the current production version, and approve a controlled rollout.

That last piece is critical. One of the reasons enterprise AI agents fail in production isn't launch day — it's month three, when company policies change, customer behavior shifts, or a new product gets added. Without a formal mechanism for updating agent behavior, the agent becomes stale. Presence builds that update loop in from the start.

The 75% Number: What's Behind It

OpenAI isn't just claiming this in theory. They deployed Presence at their own English-language phone support line — 1-888-GPT-0090 — and published production results.

The system handles open-ended requests, verifies callers, accesses account context, and takes approved actions. According to OpenAI, it met or exceeded benchmarks for frontline human support quality within weeks of launch. It now resolves 75% of inbound issues without human assistance.

Even more revealing: the Codex-powered improvement loop reduced human handoffs by 15 percentage points in just 10 days. That's not a marketing projection. That's a production deployment showing a feedback loop that actually works.

The 25% that still escalates to humans isn't a failure — it's by design. Presence is built with explicit escalation rules. Some scenarios genuinely require human judgment, empathy, or authority that an agent shouldn't have. The goal isn't 100% automation. It's reliable, auditable automation for the cases that can be handled safely, with clean handoffs for everything else.

For CIOs comparing this against internal benchmarks: most enterprise chatbot deployments achieve 30-40% containment rates in the first year. Getting to 75% typically requires years of tuning, massive QA investment, and still doesn't come with the governance infrastructure that Presence packages at launch.

Why Governance Is the Hard Part (And Why Most Teams Underestimate It)

Here's the pattern I've seen repeatedly: a team builds an AI agent for customer support. It works well in testing. They ship it. Three months later, someone in Legal updates the refund policy. Now the agent is giving incorrect answers about refunds, and nobody caught it for two weeks.

This isn't a model quality problem. It's a governance problem. And governance is fundamentally harder than model quality because it spans people, processes, and technical systems simultaneously.

Presence addresses this with several mechanisms. Before a deployment goes live, teams can test it against common requests, edge cases, and high-risk scenarios. Graders check whether the agent reached the right outcome, followed policy, used tools correctly, and escalated appropriately. Guardrails can intervene when an interaction moves outside company-defined boundaries.

After launch, production sessions, escalations, and quality signals feed back into the Codex improvement loop. The agent doesn't just run — it learns from where it's failing, surfaces proposed fixes, and lets teams approve changes before they go live.

This matters especially for regulated industries. A banking agent that gives incorrect advice about account terms isn't just a customer experience problem — it's a compliance problem. Presence's policy layer and evaluation tooling are specifically designed for environments where getting it wrong has real consequences.

Three Real Enterprise Deployments Worth Watching

OpenAI named three organizations evaluating Presence at launch. Each one tells a different story about enterprise use cases.

BBVA is exploring AI-powered voice support for everyday banking needs in Mexico. Banking is arguably the most governance-sensitive industry for AI agents: strict regulation, high stakes for individual customers, and zero tolerance for incorrect account information. The fact that a major bank is testing this in production signals serious confidence in the governance model.

SoftBank is testing natural Japanese-language customer conversations. This matters because most enterprise AI deployments are English-first and struggle badly with non-English interactions. Japanese in particular requires a completely different approach to natural language, formality levels, and cultural context. If Presence can handle production-quality Japanese customer service, it opens up enterprise deployments across every major global market.

IAG, the Australian insurance company, is exploring support during high-demand events like severe weather and natural disasters. Insurance claim peaks are exactly the scenario where human support capacity gets overwhelmed fastest. An agent that can reliably handle first-notice-of-loss calls, basic policy questions, and claims intake during a catastrophic event has enormous operational value — and enormous responsibility.

Together, these three point toward a clear pattern: Presence is being piloted in scenarios where reliability and governance matter most, not where they're nice-to-have.

The Palantir Comparison: What This Business Model Actually Means

The Forward Deployed Engineer model has been gaining attention in AI circles, but many enterprise leaders still don't fully understand what it implies for vendor relationships.

Palantir pioneered this approach — embedding engineers directly with enterprise customers to adapt its software to complex government and commercial environments. The model works for technology that is genuinely difficult to deploy and configure independently. It builds deep customer relationships, creates switching costs, and allows the vendor to learn from every deployment.

OpenAI is adopting this same model with Presence. FDEs work alongside your team to select workflows, connect internal systems, establish permissions, configure policies, test agents, and move them into production. After launch, they support ongoing evolution.

This has several practical implications for enterprise buyers.

First, Presence is not a SaaS subscription you can spin up and test next Tuesday. It requires a formal engagement with OpenAI's team, which means sales cycles, implementation timelines, and resource commitments on both sides. Plan accordingly.

Second, pricing hasn't been disclosed. That's deliberate. Enterprise software sold through FDE models typically involves custom contracts based on deployment scope, workflow complexity, and ongoing support requirements. Get your procurement team aligned early.

Third, the depth of integration creates stickiness. Once Presence is connected to your CRM, your knowledge base, your ticketing system, and your escalation workflows — and once your team has been trained on the evaluation and policy tooling — migrating to a competitor isn't a simple switch. This is worth understanding before you sign.

None of this is inherently bad. The same dynamic applies to Salesforce, SAP, and ServiceNow. But it's different from the pay-per-API-call relationship most teams have with OpenAI today.

What CIOs and CTOs Should Evaluate Now

If you're a technology leader, here are the four questions worth working through before your next board discussion.

Does your current customer-facing AI infrastructure have a formal governance layer? If your agents run on raw API calls with prompt engineering and ad-hoc testing, you have a governance gap. Presence addresses this — but so can building a proper governance layer on your existing stack. The question is whether you want to own that problem or outsource it.

What's your volume? Presence makes economic sense at scale. For low-volume workflows, the FDE engagement model may be overkill. For high-volume customer interactions — thousands of calls per day — the economics of 75% containment shift significantly.

What's your timeline? Building a comparable agent governance infrastructure in-house is a 12-18 month project for most enterprise teams, assuming you have the right ML engineering talent. If you need production-grade deployments in Q3 or Q4, the build path is probably not viable.

Are you comfortable with OpenAI having deep access to your customer interaction data? This is the question most CIOs aren't asking loudly enough. Presence requires connecting your systems to OpenAI's deployment infrastructure. What does that mean for data residency, privacy policies, and contractual protections? Get legal involved early.

What CFOs Should Model

The financial case for Presence comes down to three numbers: current support cost per contact, current agent containment rate, and volume.

The math is straightforward. If your organization handles 50,000 customer contacts per month at an average cost of $12 per contact (a typical blended rate for mixed human/automated support), your total annual support cost is $7.2 million. Moving from 40% containment to 75% containment reduces human-handled volume by ~58%, which maps roughly to a $3-4 million annual savings depending on your fixed cost structure.

That's a simplified model. Real implementations need to account for implementation costs, ongoing FDE support fees, integration expenses, and the productivity of the humans handling the escalated 25%. But the directional math is compelling — and it's the kind of P&L impact that CFOs can take to the board.

More importantly, this is recurring savings, not one-time. Unlike most technology investments that deliver efficiency gains and then plateau, agent systems that improve continuously through feedback loops compound their value over time. A containment rate that goes from 75% to 80% over 12 months compounds your savings.

The Build vs. Buy Question in 2026

A year ago, the case for building your own enterprise agent infrastructure was stronger. The models were more limited, the tooling was less mature, and the governance frameworks were largely theoretical. Building in-house meant full control and lower ongoing costs.

That calculus has shifted. Models are capable enough that the differentiation advantage of building your own is smaller. The tooling gap between what you can build and what a mature product like Presence offers is wider. And the opportunity cost of having your best ML engineers maintaining infrastructure is higher.

The case for building still makes sense in specific situations: if you have genuinely unique operational requirements that no vendor can address, if you're at massive scale where in-house economics win decisively, or if your AI team has the depth to build and maintain governance infrastructure as a core competency.

For most enterprise organizations, the build-vs-buy calculus in 2026 is shifting toward buy for foundational agent infrastructure, reserve engineering talent for differentiated applications. Presence is an early signal of where that market is heading.

The Bottom Line

OpenAI Presence represents the most significant evolution in how enterprises will deploy AI agents over the next 18 months. The 75% resolution rate is the headline, but the real story is the governance infrastructure that makes that number possible — and sustainable.

For CIOs: start evaluating this now. The FDE engagement model means there are no quick pilots. Getting Presence into production requires relationship-building and planning that starts months before the deployment date.

For CFOs: run the containment math for your top three customer-facing workflows. If volume is there, the financial case is strong and getting stronger as the product matures.

For CTO/architects: don't mistake Presence for a commodity API. This is a governed deployment product that connects deeply with your internal systems. Treat vendor evaluation with the same rigor you'd apply to a CRM or ERP selection.

The most important shift here isn't the product itself — it's what it signals. OpenAI is no longer just a model provider. With Presence, it's becoming the systems integrator for enterprise AI. That's a different kind of relationship, with different implications for enterprise architecture, vendor strategy, and long-term AI investment decisions.

The organizations piloting this now — BBVA, SoftBank, IAG — are building institutional knowledge in how to govern production AI agents at scale. That knowledge compounds. Start building yours.


For more on enterprise AI agent deployment strategy, see AI Agent Deployments That Actually Stick and The Real Cost of Enterprise AI: What Nobody Tells You.

Follow Rajesh Beri on LinkedIn and X for daily enterprise AI insights.

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

OpenAI Presence: 75% Issue Resolution, No Humans Needed

Photo by Google DeepMind on Pexels

When OpenAI announced Presence last week, the headline number was 75% — the share of inbound support calls resolved without any human assistance. That stat is real. But the bigger story is what it signals about where enterprise AI is heading, and why the window for building your own solution is getting shorter.

Most enterprise AI deployments I've seen over the past two years share a common problem: the first pilot works great. The second one requires twice the infrastructure. By the fourth or fifth deployment, you're drowning in custom prompt chains, brittle integrations, and a team of engineers who spend more time firefighting than building. The governance problem — making agents behave consistently, safely, and adaptably — has been the hard part nobody talks about at conferences.

OpenAI Presence is a direct answer to that problem. And based on their early production results, it's working.

What OpenAI Presence Actually Is

Presence is OpenAI's enterprise product for deploying governed AI agents across voice and chat workflows. It launched July 22, 2026, and is available immediately through a limited general availability program — but it's not self-serve. You need OpenAI Forward Deployed Engineers (FDEs) and select global systems integrators to set it up.

The core idea: instead of asking your engineering team to stitch together models, APIs, security controls, and evaluation tooling, OpenAI packages that entire infrastructure and deploys it alongside you.

Each Presence deployment starts with a specific job — resolve billing issues, handle insurance claims, answer IT service requests. The agent receives only the knowledge and system access needed for that job. Your organization sets the policies: what the agent can do independently, which actions need human approval, and when to escalate entirely.

The product includes what OpenAI calls the complete operational stack: policies and standard operating procedures, guardrails, approved actions, simulation testing, evaluation tools, and a Codex-powered improvement process that monitors production behavior and proposes updates. Teams review proposed changes, test them against the current production version, and approve a controlled rollout.

That last piece is critical. One of the reasons enterprise AI agents fail in production isn't launch day — it's month three, when company policies change, customer behavior shifts, or a new product gets added. Without a formal mechanism for updating agent behavior, the agent becomes stale. Presence builds that update loop in from the start.

The 75% Number: What's Behind It

OpenAI isn't just claiming this in theory. They deployed Presence at their own English-language phone support line — 1-888-GPT-0090 — and published production results.

The system handles open-ended requests, verifies callers, accesses account context, and takes approved actions. According to OpenAI, it met or exceeded benchmarks for frontline human support quality within weeks of launch. It now resolves 75% of inbound issues without human assistance.

Even more revealing: the Codex-powered improvement loop reduced human handoffs by 15 percentage points in just 10 days. That's not a marketing projection. That's a production deployment showing a feedback loop that actually works.

The 25% that still escalates to humans isn't a failure — it's by design. Presence is built with explicit escalation rules. Some scenarios genuinely require human judgment, empathy, or authority that an agent shouldn't have. The goal isn't 100% automation. It's reliable, auditable automation for the cases that can be handled safely, with clean handoffs for everything else.

For CIOs comparing this against internal benchmarks: most enterprise chatbot deployments achieve 30-40% containment rates in the first year. Getting to 75% typically requires years of tuning, massive QA investment, and still doesn't come with the governance infrastructure that Presence packages at launch.

Why Governance Is the Hard Part (And Why Most Teams Underestimate It)

Here's the pattern I've seen repeatedly: a team builds an AI agent for customer support. It works well in testing. They ship it. Three months later, someone in Legal updates the refund policy. Now the agent is giving incorrect answers about refunds, and nobody caught it for two weeks.

This isn't a model quality problem. It's a governance problem. And governance is fundamentally harder than model quality because it spans people, processes, and technical systems simultaneously.

Presence addresses this with several mechanisms. Before a deployment goes live, teams can test it against common requests, edge cases, and high-risk scenarios. Graders check whether the agent reached the right outcome, followed policy, used tools correctly, and escalated appropriately. Guardrails can intervene when an interaction moves outside company-defined boundaries.

After launch, production sessions, escalations, and quality signals feed back into the Codex improvement loop. The agent doesn't just run — it learns from where it's failing, surfaces proposed fixes, and lets teams approve changes before they go live.

This matters especially for regulated industries. A banking agent that gives incorrect advice about account terms isn't just a customer experience problem — it's a compliance problem. Presence's policy layer and evaluation tooling are specifically designed for environments where getting it wrong has real consequences.

Three Real Enterprise Deployments Worth Watching

OpenAI named three organizations evaluating Presence at launch. Each one tells a different story about enterprise use cases.

BBVA is exploring AI-powered voice support for everyday banking needs in Mexico. Banking is arguably the most governance-sensitive industry for AI agents: strict regulation, high stakes for individual customers, and zero tolerance for incorrect account information. The fact that a major bank is testing this in production signals serious confidence in the governance model.

SoftBank is testing natural Japanese-language customer conversations. This matters because most enterprise AI deployments are English-first and struggle badly with non-English interactions. Japanese in particular requires a completely different approach to natural language, formality levels, and cultural context. If Presence can handle production-quality Japanese customer service, it opens up enterprise deployments across every major global market.

IAG, the Australian insurance company, is exploring support during high-demand events like severe weather and natural disasters. Insurance claim peaks are exactly the scenario where human support capacity gets overwhelmed fastest. An agent that can reliably handle first-notice-of-loss calls, basic policy questions, and claims intake during a catastrophic event has enormous operational value — and enormous responsibility.

Together, these three point toward a clear pattern: Presence is being piloted in scenarios where reliability and governance matter most, not where they're nice-to-have.

The Palantir Comparison: What This Business Model Actually Means

The Forward Deployed Engineer model has been gaining attention in AI circles, but many enterprise leaders still don't fully understand what it implies for vendor relationships.

Palantir pioneered this approach — embedding engineers directly with enterprise customers to adapt its software to complex government and commercial environments. The model works for technology that is genuinely difficult to deploy and configure independently. It builds deep customer relationships, creates switching costs, and allows the vendor to learn from every deployment.

OpenAI is adopting this same model with Presence. FDEs work alongside your team to select workflows, connect internal systems, establish permissions, configure policies, test agents, and move them into production. After launch, they support ongoing evolution.

This has several practical implications for enterprise buyers.

First, Presence is not a SaaS subscription you can spin up and test next Tuesday. It requires a formal engagement with OpenAI's team, which means sales cycles, implementation timelines, and resource commitments on both sides. Plan accordingly.

Second, pricing hasn't been disclosed. That's deliberate. Enterprise software sold through FDE models typically involves custom contracts based on deployment scope, workflow complexity, and ongoing support requirements. Get your procurement team aligned early.

Third, the depth of integration creates stickiness. Once Presence is connected to your CRM, your knowledge base, your ticketing system, and your escalation workflows — and once your team has been trained on the evaluation and policy tooling — migrating to a competitor isn't a simple switch. This is worth understanding before you sign.

None of this is inherently bad. The same dynamic applies to Salesforce, SAP, and ServiceNow. But it's different from the pay-per-API-call relationship most teams have with OpenAI today.

What CIOs and CTOs Should Evaluate Now

If you're a technology leader, here are the four questions worth working through before your next board discussion.

Does your current customer-facing AI infrastructure have a formal governance layer? If your agents run on raw API calls with prompt engineering and ad-hoc testing, you have a governance gap. Presence addresses this — but so can building a proper governance layer on your existing stack. The question is whether you want to own that problem or outsource it.

What's your volume? Presence makes economic sense at scale. For low-volume workflows, the FDE engagement model may be overkill. For high-volume customer interactions — thousands of calls per day — the economics of 75% containment shift significantly.

What's your timeline? Building a comparable agent governance infrastructure in-house is a 12-18 month project for most enterprise teams, assuming you have the right ML engineering talent. If you need production-grade deployments in Q3 or Q4, the build path is probably not viable.

Are you comfortable with OpenAI having deep access to your customer interaction data? This is the question most CIOs aren't asking loudly enough. Presence requires connecting your systems to OpenAI's deployment infrastructure. What does that mean for data residency, privacy policies, and contractual protections? Get legal involved early.

What CFOs Should Model

The financial case for Presence comes down to three numbers: current support cost per contact, current agent containment rate, and volume.

The math is straightforward. If your organization handles 50,000 customer contacts per month at an average cost of $12 per contact (a typical blended rate for mixed human/automated support), your total annual support cost is $7.2 million. Moving from 40% containment to 75% containment reduces human-handled volume by ~58%, which maps roughly to a $3-4 million annual savings depending on your fixed cost structure.

That's a simplified model. Real implementations need to account for implementation costs, ongoing FDE support fees, integration expenses, and the productivity of the humans handling the escalated 25%. But the directional math is compelling — and it's the kind of P&L impact that CFOs can take to the board.

More importantly, this is recurring savings, not one-time. Unlike most technology investments that deliver efficiency gains and then plateau, agent systems that improve continuously through feedback loops compound their value over time. A containment rate that goes from 75% to 80% over 12 months compounds your savings.

The Build vs. Buy Question in 2026

A year ago, the case for building your own enterprise agent infrastructure was stronger. The models were more limited, the tooling was less mature, and the governance frameworks were largely theoretical. Building in-house meant full control and lower ongoing costs.

That calculus has shifted. Models are capable enough that the differentiation advantage of building your own is smaller. The tooling gap between what you can build and what a mature product like Presence offers is wider. And the opportunity cost of having your best ML engineers maintaining infrastructure is higher.

The case for building still makes sense in specific situations: if you have genuinely unique operational requirements that no vendor can address, if you're at massive scale where in-house economics win decisively, or if your AI team has the depth to build and maintain governance infrastructure as a core competency.

For most enterprise organizations, the build-vs-buy calculus in 2026 is shifting toward buy for foundational agent infrastructure, reserve engineering talent for differentiated applications. Presence is an early signal of where that market is heading.

The Bottom Line

OpenAI Presence represents the most significant evolution in how enterprises will deploy AI agents over the next 18 months. The 75% resolution rate is the headline, but the real story is the governance infrastructure that makes that number possible — and sustainable.

For CIOs: start evaluating this now. The FDE engagement model means there are no quick pilots. Getting Presence into production requires relationship-building and planning that starts months before the deployment date.

For CFOs: run the containment math for your top three customer-facing workflows. If volume is there, the financial case is strong and getting stronger as the product matures.

For CTO/architects: don't mistake Presence for a commodity API. This is a governed deployment product that connects deeply with your internal systems. Treat vendor evaluation with the same rigor you'd apply to a CRM or ERP selection.

The most important shift here isn't the product itself — it's what it signals. OpenAI is no longer just a model provider. With Presence, it's becoming the systems integrator for enterprise AI. That's a different kind of relationship, with different implications for enterprise architecture, vendor strategy, and long-term AI investment decisions.

The organizations piloting this now — BBVA, SoftBank, IAG — are building institutional knowledge in how to govern production AI agents at scale. That knowledge compounds. Start building yours.


For more on enterprise AI agent deployment strategy, see AI Agent Deployments That Actually Stick and The Real Cost of Enterprise AI: What Nobody Tells You.

Follow Rajesh Beri on LinkedIn and X for daily enterprise AI insights.

Share:
THE DAILY BRIEF
OpenAIEnterprise AIAI AgentsCustomer Service AIAgentic AI
OpenAI Presence: 75% Issue Resolution, No Humans Needed

OpenAI Presence deploys production-ready AI agents that resolve 75% of enterprise issues without human handoffs. What CIOs and CFOs need to know now.

By Rajesh Beri·July 24, 2026·11 min read

When OpenAI announced Presence last week, the headline number was 75% — the share of inbound support calls resolved without any human assistance. That stat is real. But the bigger story is what it signals about where enterprise AI is heading, and why the window for building your own solution is getting shorter.

Most enterprise AI deployments I've seen over the past two years share a common problem: the first pilot works great. The second one requires twice the infrastructure. By the fourth or fifth deployment, you're drowning in custom prompt chains, brittle integrations, and a team of engineers who spend more time firefighting than building. The governance problem — making agents behave consistently, safely, and adaptably — has been the hard part nobody talks about at conferences.

OpenAI Presence is a direct answer to that problem. And based on their early production results, it's working.

What OpenAI Presence Actually Is

Presence is OpenAI's enterprise product for deploying governed AI agents across voice and chat workflows. It launched July 22, 2026, and is available immediately through a limited general availability program — but it's not self-serve. You need OpenAI Forward Deployed Engineers (FDEs) and select global systems integrators to set it up.

The core idea: instead of asking your engineering team to stitch together models, APIs, security controls, and evaluation tooling, OpenAI packages that entire infrastructure and deploys it alongside you.

Each Presence deployment starts with a specific job — resolve billing issues, handle insurance claims, answer IT service requests. The agent receives only the knowledge and system access needed for that job. Your organization sets the policies: what the agent can do independently, which actions need human approval, and when to escalate entirely.

The product includes what OpenAI calls the complete operational stack: policies and standard operating procedures, guardrails, approved actions, simulation testing, evaluation tools, and a Codex-powered improvement process that monitors production behavior and proposes updates. Teams review proposed changes, test them against the current production version, and approve a controlled rollout.

That last piece is critical. One of the reasons enterprise AI agents fail in production isn't launch day — it's month three, when company policies change, customer behavior shifts, or a new product gets added. Without a formal mechanism for updating agent behavior, the agent becomes stale. Presence builds that update loop in from the start.

The 75% Number: What's Behind It

OpenAI isn't just claiming this in theory. They deployed Presence at their own English-language phone support line — 1-888-GPT-0090 — and published production results.

The system handles open-ended requests, verifies callers, accesses account context, and takes approved actions. According to OpenAI, it met or exceeded benchmarks for frontline human support quality within weeks of launch. It now resolves 75% of inbound issues without human assistance.

Even more revealing: the Codex-powered improvement loop reduced human handoffs by 15 percentage points in just 10 days. That's not a marketing projection. That's a production deployment showing a feedback loop that actually works.

The 25% that still escalates to humans isn't a failure — it's by design. Presence is built with explicit escalation rules. Some scenarios genuinely require human judgment, empathy, or authority that an agent shouldn't have. The goal isn't 100% automation. It's reliable, auditable automation for the cases that can be handled safely, with clean handoffs for everything else.

For CIOs comparing this against internal benchmarks: most enterprise chatbot deployments achieve 30-40% containment rates in the first year. Getting to 75% typically requires years of tuning, massive QA investment, and still doesn't come with the governance infrastructure that Presence packages at launch.

Why Governance Is the Hard Part (And Why Most Teams Underestimate It)

Here's the pattern I've seen repeatedly: a team builds an AI agent for customer support. It works well in testing. They ship it. Three months later, someone in Legal updates the refund policy. Now the agent is giving incorrect answers about refunds, and nobody caught it for two weeks.

This isn't a model quality problem. It's a governance problem. And governance is fundamentally harder than model quality because it spans people, processes, and technical systems simultaneously.

Presence addresses this with several mechanisms. Before a deployment goes live, teams can test it against common requests, edge cases, and high-risk scenarios. Graders check whether the agent reached the right outcome, followed policy, used tools correctly, and escalated appropriately. Guardrails can intervene when an interaction moves outside company-defined boundaries.

After launch, production sessions, escalations, and quality signals feed back into the Codex improvement loop. The agent doesn't just run — it learns from where it's failing, surfaces proposed fixes, and lets teams approve changes before they go live.

This matters especially for regulated industries. A banking agent that gives incorrect advice about account terms isn't just a customer experience problem — it's a compliance problem. Presence's policy layer and evaluation tooling are specifically designed for environments where getting it wrong has real consequences.

Three Real Enterprise Deployments Worth Watching

OpenAI named three organizations evaluating Presence at launch. Each one tells a different story about enterprise use cases.

BBVA is exploring AI-powered voice support for everyday banking needs in Mexico. Banking is arguably the most governance-sensitive industry for AI agents: strict regulation, high stakes for individual customers, and zero tolerance for incorrect account information. The fact that a major bank is testing this in production signals serious confidence in the governance model.

SoftBank is testing natural Japanese-language customer conversations. This matters because most enterprise AI deployments are English-first and struggle badly with non-English interactions. Japanese in particular requires a completely different approach to natural language, formality levels, and cultural context. If Presence can handle production-quality Japanese customer service, it opens up enterprise deployments across every major global market.

IAG, the Australian insurance company, is exploring support during high-demand events like severe weather and natural disasters. Insurance claim peaks are exactly the scenario where human support capacity gets overwhelmed fastest. An agent that can reliably handle first-notice-of-loss calls, basic policy questions, and claims intake during a catastrophic event has enormous operational value — and enormous responsibility.

Together, these three point toward a clear pattern: Presence is being piloted in scenarios where reliability and governance matter most, not where they're nice-to-have.

The Palantir Comparison: What This Business Model Actually Means

The Forward Deployed Engineer model has been gaining attention in AI circles, but many enterprise leaders still don't fully understand what it implies for vendor relationships.

Palantir pioneered this approach — embedding engineers directly with enterprise customers to adapt its software to complex government and commercial environments. The model works for technology that is genuinely difficult to deploy and configure independently. It builds deep customer relationships, creates switching costs, and allows the vendor to learn from every deployment.

OpenAI is adopting this same model with Presence. FDEs work alongside your team to select workflows, connect internal systems, establish permissions, configure policies, test agents, and move them into production. After launch, they support ongoing evolution.

This has several practical implications for enterprise buyers.

First, Presence is not a SaaS subscription you can spin up and test next Tuesday. It requires a formal engagement with OpenAI's team, which means sales cycles, implementation timelines, and resource commitments on both sides. Plan accordingly.

Second, pricing hasn't been disclosed. That's deliberate. Enterprise software sold through FDE models typically involves custom contracts based on deployment scope, workflow complexity, and ongoing support requirements. Get your procurement team aligned early.

Third, the depth of integration creates stickiness. Once Presence is connected to your CRM, your knowledge base, your ticketing system, and your escalation workflows — and once your team has been trained on the evaluation and policy tooling — migrating to a competitor isn't a simple switch. This is worth understanding before you sign.

None of this is inherently bad. The same dynamic applies to Salesforce, SAP, and ServiceNow. But it's different from the pay-per-API-call relationship most teams have with OpenAI today.

What CIOs and CTOs Should Evaluate Now

If you're a technology leader, here are the four questions worth working through before your next board discussion.

Does your current customer-facing AI infrastructure have a formal governance layer? If your agents run on raw API calls with prompt engineering and ad-hoc testing, you have a governance gap. Presence addresses this — but so can building a proper governance layer on your existing stack. The question is whether you want to own that problem or outsource it.

What's your volume? Presence makes economic sense at scale. For low-volume workflows, the FDE engagement model may be overkill. For high-volume customer interactions — thousands of calls per day — the economics of 75% containment shift significantly.

What's your timeline? Building a comparable agent governance infrastructure in-house is a 12-18 month project for most enterprise teams, assuming you have the right ML engineering talent. If you need production-grade deployments in Q3 or Q4, the build path is probably not viable.

Are you comfortable with OpenAI having deep access to your customer interaction data? This is the question most CIOs aren't asking loudly enough. Presence requires connecting your systems to OpenAI's deployment infrastructure. What does that mean for data residency, privacy policies, and contractual protections? Get legal involved early.

What CFOs Should Model

The financial case for Presence comes down to three numbers: current support cost per contact, current agent containment rate, and volume.

The math is straightforward. If your organization handles 50,000 customer contacts per month at an average cost of $12 per contact (a typical blended rate for mixed human/automated support), your total annual support cost is $7.2 million. Moving from 40% containment to 75% containment reduces human-handled volume by ~58%, which maps roughly to a $3-4 million annual savings depending on your fixed cost structure.

That's a simplified model. Real implementations need to account for implementation costs, ongoing FDE support fees, integration expenses, and the productivity of the humans handling the escalated 25%. But the directional math is compelling — and it's the kind of P&L impact that CFOs can take to the board.

More importantly, this is recurring savings, not one-time. Unlike most technology investments that deliver efficiency gains and then plateau, agent systems that improve continuously through feedback loops compound their value over time. A containment rate that goes from 75% to 80% over 12 months compounds your savings.

The Build vs. Buy Question in 2026

A year ago, the case for building your own enterprise agent infrastructure was stronger. The models were more limited, the tooling was less mature, and the governance frameworks were largely theoretical. Building in-house meant full control and lower ongoing costs.

That calculus has shifted. Models are capable enough that the differentiation advantage of building your own is smaller. The tooling gap between what you can build and what a mature product like Presence offers is wider. And the opportunity cost of having your best ML engineers maintaining infrastructure is higher.

The case for building still makes sense in specific situations: if you have genuinely unique operational requirements that no vendor can address, if you're at massive scale where in-house economics win decisively, or if your AI team has the depth to build and maintain governance infrastructure as a core competency.

For most enterprise organizations, the build-vs-buy calculus in 2026 is shifting toward buy for foundational agent infrastructure, reserve engineering talent for differentiated applications. Presence is an early signal of where that market is heading.

The Bottom Line

OpenAI Presence represents the most significant evolution in how enterprises will deploy AI agents over the next 18 months. The 75% resolution rate is the headline, but the real story is the governance infrastructure that makes that number possible — and sustainable.

For CIOs: start evaluating this now. The FDE engagement model means there are no quick pilots. Getting Presence into production requires relationship-building and planning that starts months before the deployment date.

For CFOs: run the containment math for your top three customer-facing workflows. If volume is there, the financial case is strong and getting stronger as the product matures.

For CTO/architects: don't mistake Presence for a commodity API. This is a governed deployment product that connects deeply with your internal systems. Treat vendor evaluation with the same rigor you'd apply to a CRM or ERP selection.

The most important shift here isn't the product itself — it's what it signals. OpenAI is no longer just a model provider. With Presence, it's becoming the systems integrator for enterprise AI. That's a different kind of relationship, with different implications for enterprise architecture, vendor strategy, and long-term AI investment decisions.

The organizations piloting this now — BBVA, SoftBank, IAG — are building institutional knowledge in how to govern production AI agents at scale. That knowledge compounds. Start building yours.


For more on enterprise AI agent deployment strategy, see AI Agent Deployments That Actually Stick and The Real Cost of Enterprise AI: What Nobody Tells You.

Follow Rajesh Beri on LinkedIn and X for daily enterprise AI insights.

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →