Every consumption-priced agent contract pays the vendor more when the agent works badly. That is not a cynical reading; it is the arithmetic. An April 2026 preprint from researchers at Michigan, Stanford and Microsoft ran eight frontier models over SWE-bench Verified and found that runs on the same task can differ by up to 30x in total tokens, with accuracy peaking at intermediate cost. Those were coding agents, not support agents, and no customer-service vendor publishes anything comparable — but it is the best-measured evidence available, and the mechanism is the meter, not the domain. On a per-token or per-action meter, you absorb that variance. On a per-resolution meter, the vendor does.
That single fact should decide most of your agentic AI purchasing, and it currently decides almost none of it.
The Verdict, and the Workload It Is Priced On
Buy outcome-based pricing for customer-facing agents, per-action credits for internal automation you control, and per-seat only where usage is genuinely bounded. Never sign a pure consumption meter without a hard cap and a defined behaviour at that cap.
Vendors quote incomparable units on purpose — conversations, credits, assists, exchanges, resolutions, vCPU-hours — so the only honest comparison is one workload priced every way. Here is the workload used throughout this piece:
Workload W: a customer-facing support agent handling 10,000 conversations a month, averaging 8 billable agent steps per conversation, of which 60% (6,000) close without a human. Fifty human support staff remain on staff.
| Model | Representative rate (checked 20 Aug 2026) | Cost on Workload W | What it pays the vendor to do | Who carries the variance |
|---|---|---|---|---|
| Per resolution | Fin: $0.99/outcome | $5,940/mo and up | Close it without a human — or finish a handoff procedure | Vendor |
| Per resolution | Salesforce Agentforce Help Agent: $2.00/resolution | $12,000/mo | Succeed without a human | Vendor |
| Per resolution + seats | Zendesk: $1.50 committed / $2.00 overage, plus $115/agent/mo | $14,750/mo | Succeed, and keep your seat count | Vendor (usage), you (seats) |
| Per action / credit | Salesforce Agentforce Flex Credits: 20 credits ($0.10)/action | $8,000/mo | Take more steps | You |
| Per action / credit | Microsoft Copilot Studio: $200 / 25,000 credits | $2,800/mo — or $9,200 with a reasoning model | Take more steps, and reason longer | You |
| Per conversation | Agentforce conversations: $2.00/conversation | $20,000/mo | Get talked to | You |
| Raw consumption | Amazon Bedrock AgentCore: $0.0895/vCPU-hour | ~$28/mo infra + all model tokens, metered separately | Run longer loops | You, entirely |
| Per seat | Google Gemini Enterprise: $30/seat/mo | $1,500/mo — but cannot price this workload at all | Hire more people | Vendor, until quota |
The spread on identical work is $5,940 to $20,000 a month among the models that can actually price it. That is not a quality difference. It is a unit difference.
What Each Pricing Model Pays the Vendor To Do
A pricing model is an incentive contract, and you should read it as one before you read the rate card. Four structures dominate the market, and each one rewards a different vendor behaviour.
Per-seat pays the vendor for your headcount. Google Gemini Enterprise is billed per seat per month, from roughly $21 for Business and $30 for Standard. The vendor's revenue rises when you hire, and falls when an agent lets you hire fewer people — which is a perverse incentive on an automation product, but a beautifully forecastable one for you.
Per-action or per-credit pays the vendor for steps taken, not problems solved. Under Agentforce Flex Credits, a standard action costs 20 credits and credits are sold at $500 per 100,000 — $0.10 an action, with voice at 30 credits ($0.15) and a ceiling of 10,000 tokens per action. Cross that ceiling and one action bills as two. A chattier agent, a longer plan, a retry loop — all revenue.
Microsoft Copilot Studio is more granular and more transparent about it. Its published rate table charges 1 credit for a classic answer, 2 for a generative answer, 5 for an agent action, 10 for tenant graph grounding, and 13 per 100 agent-flow actions. Credits cost $200 per 25,000-credit pack, or $0.008 each.
Raw consumption pays the vendor for compute time, which is the most direct version of the same inversion. Bedrock AgentCore meters $0.0895 per vCPU-hour and $0.00945 per GB-hour, with the gateway at $0.005 per 1,000 API invocations. On Workload W — 90 seconds of one vCPU and 2 GB per conversation — that is about $28 a month. The infrastructure meter is a rounding error. The model tokens it burns are billed separately, and that is the real invoice.
Outcome-based pays the vendor for success. Salesforce's Agentforce Help Agent, generally available since July 2026, charges $2 per autonomous resolution with a 1,000-resolution minimum pre-purchase, and does not charge when the issue escalates to a human or the user says they did not get their answer. Fin charges $0.99 per resolution; Zendesk $1.50 on committed volume and $2.00 pay-as-you-go; Gorgias $0.60–$1.27. This is the only structure where the vendor's margin improves when its agent gets more efficient rather than more verbose.
The Same Task Can Cost 30x. You Are Buying That Variance.
Consumption pricing asks you to forecast a number that nobody — including the model itself — can forecast. That study is the clearest evidence available, and it is worth being precise about what it measured — agentic coding runs through the OpenHands framework, not support conversations: agentic tasks consume 1,000x more tokens than code chat or code reasoning, input tokens rather than output tokens drive the bill, and frontier models predict their own token usage with correlations of only 0.39, systematically underestimating. Human expert difficulty ratings barely track actual cost either. Model choice alone swung it: Kimi-K2 and Claude Sonnet 4.5 averaged over 1.5 million more tokens than GPT-5 on the same tasks.
Read that as a procurement fact. You are being asked to commit to a budget for a quantity whose run-to-run standard deviation is enormous, whose driver is a configuration the vendor controls, and which the best available predictor gets wrong in one direction.
The consequences are already visible. Forrester's July analysis, reported by TechRepublic, lists Uber burning its annual AI budget in four months, Microsoft ending Claude Code licences after exhausting a yearly allocation, and Tesla capping AI spending at $200 per week. Flexera's 2026 State of ITAM found 59% of respondents reporting increased wasted AI spend and only 31% with accurate visibility into it. Gartner's standing prediction is that over 40% of agentic AI projects will be cancelled by the end of 2027, escalating cost first on the list of reasons.
The Copilot Studio numbers show how a single configuration choice moves the bill. On Workload W with two generative answers and six agent actions per conversation, you spend 34 credits — 340,000 a month, 14 packs, $2,800. Turn on a reasoning model and Microsoft bills a second meter: the premium text-and-generative-AI-tools rate of 10 credits per 1,000 tokens, on top of the feature rate. At 8,000 tokens of reasoning per conversation that is 80 extra credits, 114 total, 46 packs — $9,200. Same agent, same traffic, 3.3x the invoice, from a dropdown.
Outcome Pricing Wins Until You Read the Definition of Resolved
Outcome pricing is the right default for customer-facing agents, and its entire risk sits in one contractual definition that the vendor currently writes.
The definitions in market are real and they differ. Zendesk's model is status-based: as soon as a conversation is escalated to a human it can no longer count as an automated resolution, and conversations flagged resolved are verified by an LLM. Salesforce bundles everything inside a 10-minute window on a call into a single billable resolution regardless of how many questions were answered, and charges nothing when the user escalates or states they did not get an answer. Fin is looser than its marketing implies. It bills a confirmed resolution and an assumed one — a customer who leaves after Fin's answer without asking for more — at the same $0.99, and it bills a second outcome type entirely: a completed procedure handoff, when Fin runs a workflow you configured to end at a human. It charges nothing when the customer asks for a human, when a procedure fails, or when it only trades greetings.
Zendesk layers a quiet period on top of that: on email and web form conversations it counts a resolution after 72 hours of inactivity, with a shorter, configurable window on messaging. That is a proxy for success, not success: a customer who gives up and calls a competitor is billed as a resolution. And it is not a Zendesk quirk — Fin's assumed resolution draws exactly the same inference. What separates them is the remedy. Fin reverses the charge if that conversation is reopened, even across billing periods. The reversal, not the definition, is the clause worth fighting for.
The market knows this. Decagon, which offers both per-conversation and per-resolution, acknowledges plainly that "defining what a resolution is can be tricky, as not all cases end wrapped up in a bow" — and reports that the majority of its customers pick per-conversation precisely to avoid the argument. Sierra and Decagon both negotiate the definition deal by deal, and neither publishes a rate.
That is the honest trade. Outcome pricing moves the token variance onto the vendor's balance sheet, and moves a measurement dispute onto yours. The measurement dispute is the better problem: it is finite, auditable, and you can write the clause. The variance is not.
Per-Seat Cannot Price an Agent That Has No Seat
Per-seat pricing is the only genuinely forecastable model on the market, and it structurally cannot price autonomous work. An agent that runs on a trigger at 3am has no seat, no user, and no licence to consume.
What vendors actually ship is a hybrid, and the hybrid is where buyers get surprised. Gemini Enterprise pairs its per-seat fee with pooled quotas — 160 assistant queries per day per seat on Standard, 30 GiB of storage per seat, drawn from a shared pool rather than capped per user — with overages billed separately, storage at $5 per GiB per month. Microsoft does the same in reverse: a Microsoft 365 Copilot licence at $30 per user per month zero-rates classic answers, generative answers and graph grounding for employee-facing agents, while everything customer-facing runs on the credit meter.
So the seat price is a floor, not a price. The question to ask any per-seat vendor is: what happens on the day my agents exceed the pooled quota? Google restricts usage unless overages are explicitly enabled on an invoiced billing account. Microsoft's answer is documented and worth reading twice: enforcement triggers at 125% of prepaid capacity, at which point custom agents are disabled and users are told "This agent is currently unavailable. It has reached its usage limit." Unused credits do not carry over month to month.
Both halves of that matter. A forecast that comes in low costs you a service outage in front of customers. A forecast that comes in high is simply forfeited. You are penalised in both directions for an estimate the vendor's own documentation ships an estimator tool to help you guess.
Where Consumption Pricing Is Actually the Right Answer
Buy consumption when you control the loop. The objection to consumption pricing is a principal-agent problem, and it disappears the moment the principal and the agent are the same party.
If your engineers write the orchestration, choose the model, set the step budget and own the retry policy, then a per-vCPU-hour or per-token meter is the fairest possible deal: you pay for what you chose to spend, and every efficiency you win is yours. That is Amazon Bedrock AgentCore, Google Vertex AI, and self-hosted frameworks. It is also why the infrastructure meter looks so cheap on Workload W — $28 against $12,000 — because it is only selling you the container, not the judgement inside it.
Consumption becomes indefensible when the vendor controls the loop and bills you for its length. UiPath sits interestingly in the middle: its conversational agent licensing bills per exchange — the user's prompt, the agent's responses, and any tool calls in between, all one unit — at 1 Agent Unit under Flex or 0.2 Platform Units under Unified pricing, with document tools at 1 AU per file or page. Bundling the tool calls into the exchange is the right design: it stops a chattier agent from being a more expensive one. What UiPath does not publish is the dollar value of an Agent Unit, which puts it in the same category as ServiceNow and Cognigy — a defensible unit at an undisclosed rate.
Seven Clauses That Decide What You Actually Pay
The rate card is the least negotiable and least important part of the contract. These seven clauses move more money.
- A hard monthly cap, and the defined behaviour at the cap. Microsoft's is 125% then agents are disabled mid-service. Get yours in writing, and get it to throttle or degrade rather than fail closed on a customer-facing agent.
- Rollover of unused prepaid units. "Credits do not carry over" is a tax on forecasting error in the direction you cannot control. Ask for quarterly or annual true-up instead of monthly expiry.
- The token ceiling inside the billable unit. Agentforce bills one action per 10,000 tokens; a 15,000-token action is two actions. That multiplier belongs in the contract, not only in the docs, because the docs can change.
- The right to switch models without repricing. Model choice is worth 1.5 million tokens per task in the SWE-bench data. If the vendor picks the model and you pay per token, the vendor is choosing your bill.
- Who pays for failure. When the agent loops, retries, or gives up, who eats those tokens? Outcome pricing answers this by construction. Every other model needs the clause.
- The outcome definition, plus an audit right and a credit-reversal process. Specify the quiet period, specify that a reopened ticket reverses the charge, and reserve the right to sample and dispute. Vendors offering outcome pricing expect this ask.
- Meter-definition stability for the term. Microsoft renamed "messages" to "Copilot Credits" on 1 September 2025 without changing the pack size or the rate. That was benign. Bind the vendor to it being benign next time.
Who Should Not Buy Each Model
Every model has a buyer it is actively wrong for, and vendors will not tell you which one you are.
- Do not buy per-seat if your agents run autonomously on triggers, or if the business case is headcount reduction. You will pay for seats the agent made unnecessary, and the pooled quota will not cover the machine traffic anyway.
- Do not buy per-action or per-credit if the vendor controls the agent's plan and you cannot see the step count before you sign. Ask for 30 days of production telemetry from a comparable customer, or run a metered pilot. Without one, your forecast is fiction.
- Do not buy per-conversation at all, on current market rates. Agentforce's $2-per-conversation model prices Workload W at $20,000 a month for work Fin bills from $5,940 — you pay full freight for every greeting, misroute and wrong-number. Salesforce still sells it, but has since added Flex Credits and then pay-per-resolution alongside it. That is the loser in this comparison.
- Do not buy raw consumption from a vendor whose agent you did not write. It is the cleanest deal in the market when you own the loop and the worst one when you do not.
- Do not buy outcome-based pricing if your work has no countable outcome. Research agents, drafting assistants and analysis copilots produce judgement, not resolutions, and forcing an outcome metric onto them creates a metric both sides will game.
What Would Change This Answer
Three things flip the recommendation, and one of them is already moving.
The first is volume. Below roughly 1,000 conversations a month, every model is noise against integration cost and outcome pricing's minimums start to bite — Salesforce requires a 1,000-resolution pre-purchase. Above about 50,000, the per-resolution premium over per-action becomes real money and a committed-volume consumption deal with a cap can win.
The second is token deflation. Outcome pricing is a hedge against token cost, and hedges get cheaper to skip when the underlying falls. If inference prices keep dropping, the vendor's margin on a fixed $2 resolution widens and the arbitrage moves back to you — which is exactly when to renegotiate rather than renew.
The third is your own observability. The reason most buyers cannot evaluate a per-action rate is that they cannot count their own actions. Instrument first. A month of real step-count and token data turns every one of these decisions from a judgement call into arithmetic, and it is the cheapest thing on this list.
The Bottom Line
Software pricing has been through this before. Per-seat licensing was invented because a seat was a good proxy for value delivered, and it held for thirty years because software did the same amount of work every time you ran it. Agents do not. The unit broke, and the industry is now running a live experiment on who absorbs the variance that broke it.
Outcome pricing is the honest answer to that question, and it is winning for exactly that reason — Salesforce, Zendesk, Fin, Sierra and Decagon have all landed on some version of it. It is not generous. It is vendors discovering they can price risk better than their customers can, and charging a premium for carrying it. Pay the premium anyway on customer-facing work; the alternative is a budget with a 30x tail.
Do not buy a unit you cannot count. And never buy one that pays somebody else for your agent's mistakes.
Continue Reading
AI Observability Pricing: Same 10 GB, $49 or $930 Vector Database Pricing: Only pgvector Publishes a Rate Best AI Coding Assistant at 500 Seats: Buy Copilot Business What RAG Actually Costs: $1,308 a Month at 10M Tokens/Day Copilot Cowork Is Billing Now: What Each Task Costs Best LLM Gateways for Cost Control: Self-Host First Your AI Router Is Trading a 10x Discount for a 2.5x One Anthropic Bid $7B for Cheaper Inference. Don't Fix Your Rate.
