The per-minute rate on a voice agent quote usually covers one layer of a four-layer stack, so two quotes that look $0.04 apart can land $2,000 a month apart. A "stack minute" is everything a live call consumes: speech-to-text (STT), the language model (LLM), text-to-speech (TTS) and the phone line. Vapi advertises $0.05 a minute for hosting and then bills each of the other layers separately, which puts a realistic configuration between $0.09 and $0.14. For an inbound contact centre with hold time and transfers, ElevenLabs Agents produced the lowest managed-platform bill on the workload below, Retell AI came second, and Bland AI cost the most. A self-assembled stack on Pipecat Cloud beat all of them, if you have the engineers to run it.
All prices were checked on each vendor's live pricing page on October 2, 2026.
The workload every number below uses: 10,000 inbound US calls a month, 4 minutes of AI-handled time per call (of which 45 seconds is silence or hold while someone looks up an account), 20% of calls transferred to a human with a 3-minute transfer leg, 25 calls at peak. That is 40,000 AI minutes and 6,000 transfer minutes a month.
| Option | What the headline rate covers | All-in per AI minute | Monthly bill at the workload | Pick it if | Skip it if |
|---|---|---|---|---|---|
| ElevenLabs Agents | Platform, STT, TTS; LLM and telephony extra | $0.082 to $0.119 effective | $3,413 to $4,913 + plan | Your calls carry hold time | You need one flat number on the invoice |
| Retell AI | Itemised; 20 concurrent calls free | $0.109 | $4,580 | You want the simplest concurrency terms | Your system prompt runs past 4,000 tokens |
| Vapi | Hosting only | $0.090 to $0.143 | $3,926 to $6,030 | Your team wants to swap every layer | Nobody owns four vendor bills |
| Bland AI | LLM, STT, TTS bundled | $0.12 + $299/month | $5,339 | One flat number is worth up to ~$1,900/month | You are price-shopping at volume |
| Pipecat Cloud (self-assembled) | Agent hosting only | $0.051 to $0.089 | $2,452 to $3,952 + engineers | You have two engineers to own it | You have none |
Low end assumes an OpenAI model at the bottom of the range Vapi lists ($0.0077/min); high end assumes the top ($0.0452/min). Retell is priced with GPT 5.4 mini and Retell's own voices.
What Does a Voice AI Stack Minute Actually Include?
A stack minute is four separately metered services, and most quotes price only one of them. Vapi's pricing page is the most honest about this because it lists every line: hosting at $0.05, Deepgram transcription at $0.0095 to $0.0099, an OpenAI model at $0.0077 to $0.0452, ElevenLabs voices at $0.0146 to $0.0238, and telephony that is free on Vapi's own numbers or $0.008 to $0.014 on Twilio. Add the cheapest of each and you get $0.0818 on a Vapi number, or $0.0903 on a Twilio inbound line at Twilio's $0.0085 rate. Add the most expensive and you get $0.1429.
Retell AI itemises the same way with different names. Its voice infrastructure is $0.055 a minute, its own voices $0.015, ElevenLabs voices through Retell $0.040, US telephony $0.015, and the LLM is a menu: GPT 5.4 mini at $0.024, Claude 5 Sonnet at $0.064, GPT 5.5 at $0.16 and GPT 6 Astra at $0.32. The model choice alone moves Retell's all-in rate from about $0.11 to above $0.40.
ElevenLabs Agents charges $0.08 a minute on every tier and states that "LLM usage is billed separately on top, based on the model you choose." ElevenLabs adds no telephony fee of its own, but your carrier (Twilio or a SIP trunk) bills separately, which makes telephony a third meter.
Bland AI is the one bundled rate in the group: $0.14 a minute on its free Start plan and $0.12 on Build, which carries a $299 monthly platform fee. Bland's pricing FAQ says the rate "covers the LLM, STT, and TTS in one number, with no token charges and no model-provider pass-throughs."
The trade is predictability against price. Bland's number never moves when you change models, because you are not choosing the model on a per-token basis. Everyone else's does.
Do You Pay for Silence, Hold and Transfer Time?
On most platforms you pay full rate for silence and hold; ElevenLabs is the exception, and transfers are where the rest diverge. This is the line item that separates an inbound service desk from an outbound sales dialler, because inbound calls are full of dead air: callers digging out an account number, the agent waiting on a CRM lookup, music while a tool call returns.
ElevenLabs bills silent periods at 5% of the usual per-minute rate. On our workload, 45 seconds of each 4-minute call drops from $0.06 to $0.003, and the ElevenLabs layer falls from $3,200 to $2,630 a month. No other platform here publishes a silence discount. Vapi's hosting fee is a flat per-minute charge, and Deepgram's own Voice Agent API is billed "based on websocket connection time", which is the honest description of how streaming voice is metered almost everywhere: the meter runs while the socket is open.
Transfers split the vendors three ways:
- Retell stops its agent fee at the handoff. The FAQ on Retell's pricing page says the voice agent fee ends once a call is transferred and only telephony continues. At $0.015 a minute on two legs, our 6,000 transfer minutes cost about $180.
- Bland bills transfer time as its own rate. Its billing docs list transfer minutes at $0.05 on Start and $0.04 on Build when you use Bland's numbers, and free when you bring your own Twilio. That is $240 a month at Build rates.
- Pipecat Cloud charges per event. SIP Refer transfers cost $0.20 each, so 2,000 transfers are $400 regardless of how long the human takes to answer.
Bland also charges $0.015 for every outbound attempt and every failed call on its telephony, with call time "prorated to the exact second." That minimum matters for an outbound campaign that dials ten numbers to reach one person. For inbound, it barely registers.
The Retell Clause That Can Double Your Bill
Retell scales billed duration by prompt size once your agent's prompt passes 4,000 LLM tokens. Its billing exceptions page gives the formula: scaling factor equals prompt tokens divided by 4,000, and the original duration is multiplied by that factor, rounded up. An 8,000-token prompt is a factor of 2.
Contact-centre prompts grow. A returns policy, an escalation matrix, a list of forbidden phrases from legal and a dozen tool descriptions will push a prompt past 4,000 tokens without anyone noticing. If the scaling applies to the full billed duration, our $4,360 of Retell AI minutes becomes $8,720, and Retell falls from second place to last. The page does not say which line items the scaled duration applies to. Get that answer in writing before you sign, and measure your prompt with a tokenizer, not by eye.
The same page sets a 10-second minimum on calls under 10 seconds where the agent speaks first with a dynamic opening message. That one is small.
What Happens When You Hit the Concurrency Limit?
Concurrency is a hard ceiling on every platform, and the overage terms differ more than the per-minute rates do. Concurrency is the number of calls your account can carry at the same moment. Size it to your peak, not your average, because the call that arrives when you are full is either rejected or billed at a premium.
- Vapi includes 4 concurrent calls with no package, 10 on the $29 Core plan, 30 on Pro, and sells extra lines at $10 each per month. When every slot is full, Vapi's documentation says "you cannot start an outbound call or accept an inbound call until a slot becomes available". Our 25-call peak needs Core plus 15 lines: $179 a month.
- Retell includes 20 concurrent calls free and charges $8 per extra slot per month, so our peak costs $40. Over the limit, Retell's concurrency docs say inbound calls wait about 40 seconds for a slot, then go to a fallback number if you have configured one or end if you have not. Optional burst capacity lets calls through above the limit with a $0.10 a minute surcharge "applied to the entire call duration," which nearly doubles the cost of every burst call.
- ElevenLabs ties concurrency to plan tier (4 on Free up to 40 on Business) and offers burst up to 3 times your limit, with excess calls billed at $0.16 a minute, double the standard rate. Our 25-call peak needs the Scale tier at 30.
- Bland gives Build accounts 50 concurrent calls and a cap of 2,000 calls a day. The daily cap is the one to read: a seasonal spike can hit it before concurrency.
Retell's terms are the cleanest for a contact centre: generous free slots, a waiting period instead of an instant reject, and an optional fallback number so the caller reaches somebody. Vapi's are the most expensive per slot and the least forgiving at the edge.
Why Cost per Resolved Call Beats Cost per Minute
The number to compare is total spend per call that ends without a human, because a cheap minute that fails to resolve the call buys you a human call as well. In a 2024 survey of 5,728 customers, Gartner found only 14% of service issues were fully resolved in self-service. Containment rates on vendor slides are usually higher than that because "contained" often means "hung up," a gap covered in our guide to AI contact center platforms.
A worked example with illustrative inputs. Stack A costs $0.09 a minute and resolves 55% of calls. Stack B costs $0.14 and resolves 70%. Each call takes 4 minutes, and you pay an assumed $6 for every call a human has to finish (use your own fully loaded figure).
| Per 1,000 calls | Stack A ($0.09/min, 55%) | Stack B ($0.14/min, 70%) |
|---|---|---|
| AI cost | $360 | $560 |
| Calls a human finishes | 450 | 300 |
| Human cost at $6 | $2,700 | $1,800 |
| Total | $3,060 | $2,360 |
Stack B costs 56% more per minute and $700 less per thousand calls. A 15-point resolution gap swamps any per-minute difference in this table. That is why the model line on a Retell or Vapi quote deserves more scrutiny than the platform fee: the model decides resolution, and resolution decides the bill. Our eleven-stack voice agent benchmark found no single stack won on a scripted bank call, which is the argument for running your own pilot on your own calls.
Where Self-Hosting One Layer Pays
Self-host text-to-speech first, transcription last, and only once you clear the break-even volume for that layer. You do not have to move the whole stack. Pipecat Cloud hosts the orchestration at $0.01 a minute for a small agent and $0.018 a minute for PSTN, with STT, LLM and TTS "billed to provider." Assembled from Deepgram Nova-3 transcription at $0.0048 a minute, Deepgram Aura-2 voices at $0.030 per 1,000 characters (about $0.011 a minute if the agent speaks 40% of the time) and the same OpenAI range, the stack runs $0.051 to $0.089 a minute.
Pushing a layer onto your own GPUs depends on utilisation. Cerebrium, which sells GPU hosting and so has an interest in the answer, puts a self-hosted stack at $0.029 a minute, with an A10 serving STT at $1.44 an hour for 160 to 180 streams and TTS at $1.20 an hour for 180 conversations. It also makes the point most calculators skip: divide the instance price by your average concurrency, which is a fraction of peak.
Run that against our workload. An always-on A10 for TTS at $1.20 an hour is about $876 a month. Aura-2 at roughly $0.011 a minute costs $440 for our 40,000 minutes, so the break-even sits near 80,000 minutes a month. STT is worse: one A10 at $1.44 an hour is about $1,051 a month against $192 of Deepgram transcription, which puts break-even near 219,000 minutes. Our Deepgram vs AssemblyAI vs Whisper comparison reached the same conclusion from the accuracy side.
Telephony is the cheap layer to move. Twilio's Elastic SIP trunking runs $0.0034 a minute for inbound local calls against $0.0085 for inbound local PSTN, which saves about $200 a month at 40,000 minutes. Do it when you renegotiate, but it will not change the vendor decision.
The LLM is where self-hosting can pay soonest, because its cost per minute has the widest range in the stack ($0.0077 to $0.32 across the vendors above). It is also the layer that decides resolution, so a cheaper model that drops your resolution rate five points loses more than it saves.
Which Platform Should You Pick?
Pick ElevenLabs Agents for an inbound contact centre with real hold time, Retell AI if you want the simplest terms and can keep your prompt short, and Pipecat when you have engineers to own it.
ElevenLabs Agents wins this workload because of the silence rate and a flat $0.08 that does not change with voice choice. It loses some of that edge to the separate LLM line and a plan you must size up for concurrency. Do not pick it if finance needs one fixed rate per minute: three meters (platform, LLM, telephony) is the minimum you will reconcile.
Retell AI has the best overage terms and a 20-call concurrency allowance that covers most mid-size queues for free. It is the most transparent itemised menu in the group. Do not pick it if your prompt is long and growing, until Retell confirms in writing what the 4,000-token scaling applies to. That clause is the difference between second place and last.
Vapi is the builder's platform: every layer is swappable and visible, and its cheapest configuration is close to the cheapest managed bill here. Do not pick it if nobody on your team will own the Deepgram, OpenAI, ElevenLabs and Twilio relationships behind it, or if your peak runs above 10 calls without a budget for $10 lines.
Bland AI is the loser on price. At our workload it is $5,339 a month against $3,413 to $4,913 for ElevenLabs and $4,580 for Retell, and its bundle removes the one lever (model choice) that most affects resolution. Do not pick it if you are buying on cost at volume. Pick it only if a single per-minute number on the invoice is worth roughly $750 to $1,900 a month to you, or for a small outbound pilot on the free Start plan.
Pipecat Cloud is the cheapest stack and the most work. Do not pick it if you cannot name the two engineers who will be paged when Deepgram changes an endpoint.
Decision Criteria That Predict Regret
Most buyers who regret a voice platform priced it at average load with a short demo prompt. The criteria that change the answer:
- Silence share above 15% of call time favours ElevenLabs. Below 5% (outbound, short scripted calls) the silence discount barely matters and Retell or Vapi close the gap.
- Prompt length past 4,000 tokens changes Retell's math entirely. Measure the production prompt, with tools, not the demo.
- Peak-to-average ratio above 3 makes concurrency terms matter more than rate. Price your worst Monday, not your monthly total divided by hours.
- Above roughly 80,000 minutes a month with an in-house platform team, self-hosting TTS starts to pay. Above roughly 220,000, STT does.
- A vendor that cannot show you resolution, defined as no repeat contact within 7 days, is quoting you a minute price for a call you may pay for twice.
Watch for ownership changes as well. Inworld's purchase of Ultravox showed how fast third-party voice terms can move after a deal.
What to Do This Quarter
This Week:
- Pull 200 recorded calls and measure three numbers: average handle time, silence share and transfer rate. Every quote depends on them.
- Count the tokens in your production system prompt, including tool definitions.
This Month:
- Send each shortlisted vendor the same workload spec (calls, minutes, silence share, transfer rate, peak concurrency, prompt tokens) and ask for a fully loaded monthly figure, in writing, that names every line item.
- Ask Retell what its prompt-token scaling applies to, and ask every vendor what happens to the 26th call at a 25-call limit.
Before Q4 Close:
- Run a two-week pilot on the top two, scored on cost per resolved call with resolution measured as no repeat contact within 7 days.
- Price a Pipecat build alongside the winner, so your renewal conversation has a number to beat.
The Bottom Line
Voice agent pricing looks like early cloud pricing: the advertised unit covers one meter, and the bill is the sum of several. The fix is the same one that worked for cloud: define the workload end to end and make every vendor quote against it. Measure your silence share and prompt size before you open a pricing page, then compare quotes on cost per resolved call.
Continue Reading
- Best AI Contact Center Platforms: Containment Isn't Resolution
- Eleven Voice Agents, One Bank Call, No Clean Winner
- Deepgram vs AssemblyAI vs Whisper: If You Self-Host, Skip Whisper
- Inworld Buys Ultravox and Upgrades Only Its Own Voices
- Mercury Voice Claims 320ms and a Benchmark Win With No Numbers
- Why 50% of Voice AI Costs Vanish with Owned Infrastructure
