Cisco Personal AI Agents: 90,000 Employees, CFO Playbook

Cisco's CFO reveals how to deploy AI agents to 90,000 workers without burning budget — efficiency-first model routing is the key. Enterprise playbook inside.

By Rajesh Beri·July 6, 2026·10 min read
Share:
Cisco Personal AI Agents: 90,000 Employees, CFO Playbook

Photo by fauxels on Pexels

When Cisco's CFO told Fortune that his company is giving every one of its 90,000 employees a personalized AI agent by the end of July 2026, most headlines focused on the scale. I'm more interested in the four words he said next: "we prioritize efficiency, not frontier models." That sentence is worth more to enterprise leaders than any vendor pitch deck you'll read this year.

Update — September 1, 2026: It shipped. On August 27 Cisco announced MyAgent, built on Circuit — its internal "secure, governed, multi-model agnostic AI platform" — and began rolling it out to all ~90,000 employees. That is general availability, not a pilot. The efficiency thesis below holds, but the cost architecture is more aggressive than this piece originally described, and one thing in the sizing argument needs correcting: the cheapest tier is not a cheap model. It is no model. Secondary coverage of Cisco's orchestration layer reports that roughly 20-30% of MyAgent requests are served by classic software automation with no generative model in the path at all, and 50-60% by open-weight models Cisco self-hosts on its own GPUs — leaving only a residual share for external frontier APIs. If those numbers hold, the budget line a CIO should be defending is GPU capacity and deterministic automation coverage, not a per-seat frontier token bill. Details, caveats and the adoption risk Cisco created for itself are in the two new sections below.

This isn't just a Cisco story. It's a blueprint — the first large-scale proof point of how an enterprise actually controls AI costs while deploying agents at workforce scale. Every CIO, CTO, and CFO in a company larger than 1,000 people should be studying what Cisco is doing, because the decisions they're making right now will become industry standard.

Let me break down what's actually happening, what it costs, and what your leadership team should take away.

The Announcement: Bigger Than It Looks

Cisco began rolling out personalized AI agents to all approximately 90,000 employees at the end of July 2026, coinciding with the company's new fiscal year, and named the product publicly on August 27. CFO Mark Patterson — who has spent 26 years at Cisco and became CFO in July 2025 — described this to Fortune, three weeks before the rollout started, as "the biggest technology transition of my career."

That's not marketing language from someone who has watched Cisco navigate the transition from physical networking hardware to cloud, from on-premises infrastructure to SaaS, from security perimeters to zero-trust. Patterson knows what a generational shift looks like.

The key differentiator here is the word "personalized." These are not generic chatbots. They're not the same tool pushed to every employee with the same prompts. Each agent is designed to learn an individual employee's role, workflows, data patterns, and preferences — then act on their behalf across multi-step tasks. A network engineer's agent looks nothing like a finance analyst's agent, which looks nothing like a sales director's agent.

That level of customization changes the risk profile and the cost equation in ways most enterprise AI deployments haven't had to confront yet.

The shipped architecture makes the mechanism concrete, and it is worth understanding because it is cheaper than the obvious reading. Personalization is not delivered by a model per employee. It is delivered by composition: reporting on the launch describes more than 800 backend subagents sitting behind the single employee-facing agent, each specialized for a task — tracking a sales-forecast variance, firing an anomaly notification — with each employee's agent pulling their own context through permissioned enterprise connectors and staying isolated from everyone else's. One agent surface, 800 shared tools, per-user context and permissions. Nobody is fine-tuning 90,000 models, and if your vendor's proposal implies that they are, ask why.

The CFO's Cost Control Secret

Here's the part that most coverage missed: Patterson was explicit that Cisco will not use frontier models — the most powerful, most expensive large language models — for every task.

"It's not going to burn a whole bunch of tokens with frontier models," he told Fortune of the system. "It knows which tool is most effective and most efficient."

Instead, Cisco built an AI routing layer that dynamically selects the most appropriate and efficient resource for each request before processing it. Simple tasks go to lighter, cheaper models. Complex reasoning tasks get routed to more powerful models only when genuinely necessary.

This is the most important architectural decision in enterprise AI right now, and most companies are getting it wrong.

In conversations with CIOs over the past year, I keep hearing the same pattern: teams default to the most capable model because it's easier to justify to stakeholders ("we use the best AI") and it removes the cognitive overhead of model selection. The problem is the bill. A complex AI agent workflow doesn't just consume a few thousand tokens like a chatbot exchange — Patterson cited industry estimates suggesting it can consume hundreds of thousands to millions of tokens in a single task chain. Multiply that by 90,000 employees, multiply it by dozens of interactions per day, and you're looking at costs that make enterprise software licensing look quaint.

Cisco's solution is an intelligent dispatch layer. Think of it as an air traffic controller for AI requests: every query gets assessed, routed to the right resource, and processed at the lowest cost that still delivers the required quality. This isn't a novel concept — it's what any financially disciplined engineering leader should be building — but Cisco is the first major enterprise to publicly confirm they've built it at this scale.

The Mix Is More Aggressive Than "Cheaper Models"

When this piece was first written, "efficiency, not frontier" read as a model-tier argument: send the easy work to a small model, keep the expensive one for hard reasoning. The reported traffic split says something stronger, and it changes the sizing math.

Coverage of Cisco's orchestration layer puts the distribution at roughly 50-60% of requests on open-weight models self-hosted on Cisco-owned GPU servers, 20-30% on classic software automation with no generative model involved at all, and only a residual share going out to external foundation models — Azure OpenAI, Anthropic's Claude, Google's Gemini. Requests are dispatched on the nature of the task, the tolerated latency, and the required reliability, and the layer is not new: Cisco has been running multi-model routing across those same external providers plus its own Deep Network Model since Circuit launched in 2023, when the internal assistant ran on Azure OpenAI alone.

Treat those percentages as reported, not confirmed. They do not appear in Cisco's own announcement, and they do not appear in the Fortune interview that one outlet credits them to — secondary coverage attributes the same three numbers to two different Cisco executives. Cisco has published nothing at this granularity. That is a reason to check the shape of the argument rather than the decimal places.

The shape survives, and it inverts the usual budget conversation in two ways.

The cheapest tier is not a cheap model — it is no model. A fifth to a third of the traffic through a workforce-scale AI assistant is, on this account, deterministic automation wearing a conversational front end. Pull a report. Open a ticket. Check a status. File an expense. Those were RPA and workflow jobs before anyone had a token budget, they cost effectively nothing per execution, they don't hallucinate, and they're auditable by inspection. The agent's real job for that slice is intent classification, not generation. If your model router has no "route to plain code" branch, you are paying inference prices for if statements — and you will keep paying them at 90,000-employee volume.

The majority tier is capacity, not consumption. Self-hosted open weights on owned GPUs is a fixed-cost, capacity-planned line item with a utilization risk, not a variable per-token bill with a runaway risk. That is a different budget object, a different failure mode, and a different team. It is also the reason a per-seat frontier-token projection is the wrong model for this kind of deployment: it prices a workload Cisco appears to have engineered down to a residual. Run the projection anyway — it tells you your ceiling — but the plan it should produce is a GPU capacity forecast and an automation-coverage target, not a volume discount negotiation.

The honest caveat is that Cisco is not a typical buyer here. It manufactures the infrastructure, it has the GPU capacity and the network engineers, and it has been building this stack since 2023. A 3,000-person company deciding today does not get 50-60% self-hosted by writing it into a plan. But the 20-30% deterministic slice is available to everyone, requires no GPUs, and is the part most enterprises are skipping.

Real Results From the Finance Team

Patterson didn't just describe the theory. He shared what's already working in his own department.

MD&A at 80-90% automation. Cisco's finance team uses AI to produce 80-90% of the first draft of the Management Discussion and Analysis section in its quarterly and annual financial filings. This is not simple text generation — the MD&A requires coherent narrative explanation of financial performance, risk factors, forward-looking statements, and segment analysis. Automating even 50% of that draft is extraordinary. Getting to 80-90% means a seasoned analyst is now in editorial mode rather than composition mode.

Investor relations intelligence. Cisco built an AI-powered tool that analyzes the company's historical financial performance alongside competitors' earnings call transcripts, then anticipates questions from individual financial analysts ahead of earnings announcements. Patterson literally knows what specific analysts are likely to ask before they ask it. That is a competitive intelligence capability that would have required a team of research analysts just three years ago.

The CFO cockpit. In development: an AI-powered executive dashboard that pulls data across products, geographies, and customer segments to generate forward-looking insights and recommendations. Patterson uses his own personal AI agent to benchmark Cisco's performance against competitors across revenue growth, EPS, R&D spending, and capital allocation — on demand.

These aren't aspirational use cases. They're in production or near-production at one of the world's largest technology companies. CFOs at Fortune 500 companies watching this should be having an immediate conversation with their finance transformation teams.

What Cisco Actually Turned On

The original interview described intent. The launch describes scope, and the scope is the part worth copying into your own requirements document.

MyAgent runs supervised autonomous workflows across Outlook, Webex, Jira and SharePoint and other internal applications — an employee states an objective and the system coordinates the steps rather than answering a question and stopping. Cisco frames the governance model as a move "from human-in-the-loop to human-in-control": the employee sets intent and remains accountable for the outcome, while the agent works in the background against approved models, approved systems and approved data pathways only. Thimaya Subaiya, Cisco's EVP of operations, is explicit that this is a product strategy as much as an IT project — "built by Cisco, for all Cisco employees, we can translate this into the blueprint that helps our customers and partners implement Enterprise AI at scale." Read the announcement accordingly. It is a reference architecture with a sales motion attached.

Two details deserve more attention than they got.

"Human-in-control" is a weaker guarantee than "human-in-the-loop," and Cisco says so. In-the-loop means a person approves each consequential action. In-control means a person set the objective and owns the result. That is the correct trade at 90,000 employees — approval queues do not scale, and an assistant that interrupts constantly gets abandoned — but it moves the control point from the action to the grant. If you adopt this framing, your compensating controls are scoping, logging and reversibility, because nobody is reading the diff before it lands. That is a governance decision, not a UX one, and it should be made by name.

Persistent memory is a retention and offboarding problem nobody costed. MyAgent remembers user preferences, past interactions and context across sessions — which is what makes it useful on day 90 instead of day 1, and also what makes it a new store of employee-derived data with no obvious owner. Before you ship this, answer four questions in writing: how long agent memory is retained, whether it is in scope for legal hold and eDiscovery, what happens to it when the employee leaves, and whether a manager can see it. Every one of those has a defensible answer. None of them has a default.

The demand signal underneath all of this is real, for whatever it is worth as a vendor-reported number: Cisco says agentic interactions on Circuit grew nearly 350% quarter over quarter before MyAgent even launched. Employees were already reaching for this. That is the argument for building the governance layer now rather than after.

What This Means for CIOs and CTOs

For technical leaders, Cisco's deployment surfaces three immediate action items.

Identity and access management is your biggest risk. An AI agent that can act on behalf of an employee needs the same access that employee has — sometimes more. That creates a new attack surface. If a threat actor compromises the agent layer, they potentially have automated, high-velocity access to everything the agent can touch. Cisco's security portfolio includes IAM and zero-trust tools, which likely gives them an advantage in building agent authentication into their own stack. Most enterprises don't have that luxury. You need a policy for agent identity — what access agents get, how it's scoped, how it's audited, and what happens when an agent behaves anomalously.

On-premises isn't dead — it's the cost control lever. Cisco built significant AI infrastructure on-premises to maintain control over both operating costs and enterprise data. This runs counter to the "everything goes to the cloud" narrative of the last decade. For enterprises where data sovereignty matters (financial services, healthcare, defense, legal) or where AI inference costs are a line item the CFO is scrutinizing, on-premises inference is back on the table. The math has changed: modern GPU clusters running efficient open-source models can be cheaper per token than cloud API calls at enterprise volume.

Model routing is now an engineering competency. The ability to evaluate, benchmark, and route AI requests to appropriate models is becoming a core platform capability — not something you outsource to a single vendor. Building this intelligently requires understanding your task taxonomy: what types of requests your employees actually generate, what model quality level each type requires, and what the cost differential between model tiers looks like at your usage volume. This is new work that most engineering teams haven't done yet.

What This Means for CFOs and Business Leaders

The Cisco story has a clear message for the business side: AI at enterprise scale is affordable if you engineer it correctly, but it is expensive if you don't.

Patterson's "efficiency, not frontier" philosophy translates directly into financial discipline. Every dollar spent on AI inference that exceeds the quality threshold for a given task is waste — the enterprise equivalent of shipping every internal memo by overnight courier when standard mail would do.

The business case math for Cisco looks compelling from the outside: $2 billion in AI-related orders in FY2025, $9 billion in AI order guidance for FY2026, and approximately 53% stock price growth year-to-date in 2026. Cisco is both selling AI infrastructure and consuming it — which gives them a uniquely honest perspective on what it actually costs and what it actually delivers.

For CFOs building the business case for their own AI agent deployments, Cisco's experience suggests several things:

Start with productivity in high-documentation roles. Finance, legal, compliance, and HR all produce enormous amounts of structured documentation — the kind where AI can automate 70-90% of first drafts, freeing professionals for analysis and judgment. The time savings are measurable and the quality bar is auditable.

Model the token cost before you model the headcount savings. Most AI ROI models I've seen at large enterprises calculate the potential labor offset without modeling the AI inference cost. At chatbot scale, this omission is tolerable. At agentic scale, it can turn a positive ROI story into an embarrassing board presentation. Build the token cost projection in from day one.

Governance investment is not optional. The risk of an AI agent making a consequential mistake — producing an incorrect financial disclosure, routing a sensitive communication to the wrong party, taking an unauthorized action in a business system — is real. The governance layer (audit trails, human review checkpoints, rollback mechanisms) costs money and slows deployment. Budget for it explicitly.

The Competitive Pressure This Creates

Cisco's announcement doesn't exist in a vacuum. Microsoft is embedding agents into Microsoft 365. Salesforce has Einstein AI agents across Sales Cloud, Service Cloud, and Marketing Cloud. ServiceNow's Now Assist operates across IT, HR, and customer service workflows. SAP and Workday are both weaving agents into their ERP and HCM platforms.

The pattern is consistent: every major enterprise software vendor is moving to a model where AI agents are the default interface for knowledge work, not an optional add-on. Cisco's 90,000-employee rollout is significant because it demonstrates the deployment model — not just the product vision.

What Cisco is proving is that this is operationally viable at scale when you engineer the cost controls correctly. That removes the last credible objection to enterprise-wide deployment: "we can't afford it at scale."

That objection is going away. Fast.

The Sequencing Problem You Inherit With the Blueprint

Here is the part of the blueprint no one wants to copy, and the most likely reason a technically sound rollout stalls.

Ten weeks before Cisco handed every remaining employee a personal AI agent, it cut nearly 4,000 jobs — about 5% of the workforce — announced May 14 alongside record quarterly revenue, with the company citing a changed cost structure and investment in AI and security. CEO Chuck Robbins named investment "in our employees' use of AI across the company" in the same breath. Cisco was not in distress; revenue was up 12% year over year. The cuts were a reallocation.

Steel-man it, because the steel-man is strong: reallocating headcount budget into AI capability is exactly what a board should expect, doing it from strength is better than doing it from weakness, and delaying the tooling would not have saved a single one of those 4,000 jobs. The two decisions are genuinely independent.

They are not independent to the person receiving the tool. An employee who watched 4,000 colleagues leave in May and receives an agent in August has a rational, unfalsifiable hypothesis about what the agent is for — and adoption of a workforce assistant is voluntary in every way that matters. You cannot mandate that someone delegate their judgment to a system they believe is auditioning for their job. This is where efficiency-first architecture stops being the binding constraint: your routing layer can be perfect and your utilization can still be 15%.

The controllable variable is not the layoff. It is the gap and the story. If a reduction and an agent rollout are going to land in the same year, decide deliberately which comes first, put real distance between them, and be specific in writing about what happens to the capacity the agent frees — because "we'll see" is heard as an answer, just not the one you meant. Cisco's own 350% quarter-over-quarter growth in agentic interactions is the number to watch here. If it holds through the rollout, the sequencing did not bite. If it flattens at general availability, it did.

Five Steps Enterprise Leaders Should Take Now

Step 1: Map your task taxonomy — and mark the ones that need no model. Before you can build an efficient routing layer, you need to know what your employees actually do. Catalog the top 20-30 task types across your largest departments. Categorize by complexity, data sensitivity, and quality requirements. Then run a second pass with one question: which of these is a deterministic workflow that only looks conversational? On Cisco's reported mix that is a fifth to a third of all traffic, and it is the cheapest, most auditable tier you will ever build. Size it first.

Step 2: Run a token cost projection. Take your current AI usage (if any), extrapolate to full-workforce deployment, and model what happens when simple chatbot exchanges become multi-step agent workflows consuming 10x-100x more tokens. The number will be uncomfortable. That discomfort is valuable — it forces the right architectural conversations early.

Step 3: Define agent identity policy. Work with your security and legal teams to define how AI agents authenticate, what access they get, and how agent activity is logged and audited. This policy should exist before you deploy a single agent, not after your first security incident.

Step 4: Start in finance or legal. Patterson's examples from the Cisco finance team are instructive. High-documentation roles with clear quality standards and measurable output are the right starting point. The ROI is visible, the risk is contained, and the learnings transfer across the organization.

Step 5: Build the cockpit before the fleet. Before deploying agents to thousands of employees, build the executive monitoring layer — the dashboard that shows agent activity, cost per user, quality metrics, and anomaly alerts. You cannot govern what you cannot see.

Bottom Line

Cisco's announcement is the most operationally instructive enterprise AI story of 2026 so far — not because of the scale, but because the CFO talked openly about the cost architecture. That candor is rare and valuable.

The core lesson: AI agents at enterprise scale are affordable when you treat model selection as a financial discipline, not a technical convenience. The enterprises that default to frontier models for every task will face budget crises. The enterprises that build intelligent routing layers will outcompete them on cost structure while delivering equivalent — sometimes better — output quality.

The shipped version sharpens that lesson rather than softening it. If the reported mix is even directionally right, Cisco got to workforce scale by moving the majority of its traffic off frontier APIs entirely — half to six-tenths onto self-hosted open weights, a fifth to a third onto code that was never a model in the first place. "Efficiency, not frontier" was not a preference for smaller models. It was a decision to stop buying tokens wherever tokens weren't the point.

MyAgent is live to 90,000 people. The next milestone worth watching is no longer the rollout — it is whether Cisco's AI order guidance moves once customers can see a real proof point at scale, and whether the 350% quarterly growth in agentic interactions survives contact with a workforce that watched 4,000 colleagues leave in May. Nine billion dollars in AI orders for FY2026 is already a strong signal. Usage after general availability is the honest one.

Route the easy work to code. Route the bulk to hardware you own. Save the frontier model for the problems that actually deserve one.


Sources: Fortune interview with Cisco CFO Mark Patterson (July 1, 2026); Cisco earnings disclosures; Cisco Newsroom blog, "MyAgent and the Rise of Ambient Intelligence" (August 27, 2026); PYMNTS, TechCrunch and secondary trade coverage of the MyAgent rollout

Continue Reading

Share:

Frequently Asked Questions

How is Cisco keeping AI agent costs down across 90,000 employees?

Cisco built an AI routing layer that dynamically assigns each request to the most efficient model rather than defaulting to expensive frontier models. As CFO Mark Patterson put it, "We're not going to burn a whole bunch of tokens with frontier models." Simple tasks go to lighter, cheaper models; only genuinely complex reasoning is routed to more powerful ones. Much of the infrastructure also runs on-premises to control operating costs and data security.

When is Cisco rolling out personalized AI agents to all employees?

Cisco is deploying a personalized AI agent to each of its roughly 90,000 employees starting at the end of July 2026, timed to the start of its new fiscal year. Each agent is designed to learn an individual employee's role, workflows, and data patterns and act on their behalf across multi-step tasks.

What has Cisco's finance team already automated with AI?

Cisco's finance team now uses AI to produce 80-90% of the first draft of the Management Discussion and Analysis (MD&A) section of its filings. It also built an investor-relations tool that analyzes historical performance against competitors' earnings calls to anticipate analyst questions, and is developing a "CFO cockpit" dashboard that synthesizes performance data across products, geographies, and segments.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →