Self-hosting speech-to-text starts paying somewhere between 2 and 6 million call minutes a month, and at that scale the model you should self-host is not Whisper. On role-played agent-customer calls across 14 English accents, Whisper large-v3 scored 15.0% word error rate against 9.2% for NVIDIA's Parakeet. On the same A100, Parakeet also transcribes long audio about 14 times faster. Whisper is the open model everyone knows. It is the wrong one to run in a contact centre.
The verdict: buy AssemblyAI Universal-3.5 Pro for post-call transcription. It has the best independent accuracy of the three, the cheapest published price once diarisation and PII redaction are added, and a free opt-out from model training. Buy Deepgram Nova-3 and its Flux turn-taking model only where a live voice agent needs fast turn detection, and only after legal has read its training terms. Self-host NVIDIA Parakeet TDT 0.6B v3 once you pass the crossover below and have two engineers to own it. OpenAI GPT-Transcribe makes sense for teams that already have OpenAI zero-data-retention approved. Self-hosted OpenAI Whisper large-v3 comes last.
The workload everything here is normalised to: 5 million call minutes a month (about 83,300 call-hours), English with non-native accents, recorded as split-channel stereo the way Amazon Connect stores agent audio in the right channel and customer audio in the left. The job is post-call transcription with speaker attribution and PII text redaction. List prices checked 2026-09-22.
| Option | Batch list price | Live list price | Independent WER | Diarisation | PII redaction | Trains on your audio? | Verdict |
|---|---|---|---|---|---|---|---|
| AssemblyAI Universal-3.5 Pro | $0.21/hr | $0.45/hr (Universal-Streaming $0.15/hr) | 3.1% AA-WER* | +$0.02/hr batch | +$0.08/hr, 50 languages | Some files, after redaction; free opt-out | Default for post-call |
| Deepgram Nova-3 (+ Flux) | $0.0043/min | $0.0048/min promo, $0.0077 regular | 5.2% AA-WER | Included in batch | +$0.0020/min; entities English-only | Yes, by default | Live voice agents |
| OpenAI GPT-Transcribe | $0.0045/min | $0.017/min | 3.3% AA-WER | Separate model, $0.006/min | None | No; ZDR-eligible | Existing OpenAI shops |
| NVIDIA Parakeet TDT 0.6B v3 | ~$0.0025 GPU per audio-hour† | Build it | 9.2% on accented calls | Bring your own | Bring your own | Never leaves your VPC | The self-host pick |
| OpenAI Whisper large-v3 | ~$0.036 GPU per audio-hour† | ~3.3 s lag | 15.0% on accented calls | Bring your own | Bring your own | Never leaves your VPC | The loser |
*Measured on Universal-3 Pro, the predecessor. †An A100 80GB at Modal's per-second rate, at full load, using the Open ASR long-form throughput figures. Sources are in the sections below.
Which speech-to-text model is most accurate on accented, noisy call audio?
AssemblyAI leads the independent benchmarks, and Deepgram trails on word error rate. None of the public benchmarks tests what you actually record: 8 kHz phone audio with accents and crosstalk. Treat every number below as a way to build a shortlist. None of them settles the decision.
The most relevant independent index is Artificial Analysis's AA-WER. Half of it is weighted on AA-AgentTalk, a proprietary set of speech from voice-agent use cases, and a quarter on Earnings22, which is earnings calls with technical language and overlapping speakers. On that index, as of 2026-09-22, AssemblyAI Universal-3 Pro scores 3.1%, OpenAI GPT Transcribe 3.3%, Whisper Large v3 4.1% and Deepgram Nova-3 5.2%. One caveat applies. The AssemblyAI model on that board is the predecessor to the current Universal-3.5 Pro, which AssemblyAI released on 7 July 2026 and whose own numbers are vendor-run. ElevenLabs Scribe v2 sits above all of them at 2.2%, which is worth a look if you are shortlisting outside these three.
Hugging Face's Open ASR Leaderboard shows the same order on long recordings. On the long-form English track, AssemblyAI Universal 3 Pro averages 8.34% WER, Parakeet TDT 0.6B v3 10.7% and Whisper Large v3 11.2%. Deepgram is not in that table.
The benchmark closest to your floor is AppTek's call-centre corpus: 128.6 hours of role-played agent-customer conversations across fourteen English accents and sixteen service scenarios. It only tested open models, and it is 16 kHz wideband rather than telephony. Inside those limits the result is blunt. With Silero segmentation, Whisper Large v3 scored 15.0% WER against 9.2% for Parakeet v3 and 8.3% for Qwen3-ASR 1.7B. Whisper's worst accents were Chinese-accented English at 20.1%, Singaporean at 18.0% and Scottish at 17.3%. The authors' summary is worth giving to anyone on your team who picked a model from a US-English leaderboard: "good performance on general American English benchmarks does not necessarily generalize to other accents."
Now the steel-man for Whisper. On AA-WER, Whisper large-v3 (4.1%) beats Deepgram Nova-3 (5.2%). It also beats Parakeet v2 (6.4%) there, so Whisper is not a bad model. Its problem is fragility. In the same AppTek study, fixed-length chunking pushed Whisper Large v2 to 48.4% WER. Segmentation is exactly the part you own when you self-host, and it decides whether Whisper is a 4% model or a 48% one.
Noise hurts everything. Parakeet v3's own model card shows WER rising from 6.34% on clean audio to 19.88% at −5 dB SNR. Aggregate WER also hides the errors your QA team cares about. On accented conversational English from India, Indonesia and Latin America, a 110M-parameter Parakeet production model recalled only 53.6% to 55.2% of named entities. Those entities are account numbers, product names and people's names, which is everything a compliance reviewer searches for.
Vendors publish sharper claims than any of this. AssemblyAI says Universal-3.5 Pro Realtime scored 6.99% WER on real agent conversations against 15.58% for Deepgram Flux. That is AssemblyAI's own test on AssemblyAI's own dataset. Put it on the shortlist and nowhere else.
Whisper hallucinates in silence, and call recordings are mostly silence
Whisper invents text when the audio goes quiet, and a split-channel call recording goes quiet on each channel for roughly half the call. It is a documented failure mode of the model, and contact-centre audio triggers it more than almost any other kind.
Cornell's "Careless Whisper" study, testing OpenAI's hosted Whisper API in spring 2023, found that roughly 1% of transcriptions contained entire hallucinated phrases or sentences, 38% of which included explicit harms. Hallucinations were concentrated in speakers "with longer shares of non-vocal durations." A December 2023 retest found many of those examples fixed, but not all. OpenAI's own model card for large-v3 warns that predictions "may include texts that are not actually spoken in the audio input" and that the model tends toward "repetitive texts."
Now think about your recordings. In a stereo recording, the customer channel is silent every time the agent speaks, and the reverse. Hold time adds more, and Amazon Connect notes that when a customer is on hold, the agent is still recorded. If you feed that to Whisper raw, you are feeding it the exact condition that produces fabricated sentences. Those fabricated sentences then land in a transcript that your QA, compliance and dispute teams treat as the record.
It can be engineered around. The usual fix in the long-running openai/whisper discussion on hallucination is voice-activity detection, with Silero-VAD placed in front of the model. Once you add VAD, though, you are maintaining a pipeline. The model is the smallest part of it.
Why real-time transcription is a different purchase from batch
Live transcription is a latency contract, and Whisper cannot meet it. For agent assist the vendors are roughly at parity. For voice agents, Deepgram's Flux was built for the job.
Whisper has a 30-second receptive field. It was built to process chunks, not streams. The best-known research effort to stream it, Whisper-Streaming, reports 3.3 seconds of latency on unsegmented long-form speech. That is fine for captions. It is too slow for a suggestion that has to reach an agent before the customer finishes a sentence, and a voice agent that pauses that long gets talked over. Parakeet v3 supports streaming through a dedicated inference script, but the serving layer, the end-of-turn logic and the capacity planning are all yours to build.
Deepgram's Flux combines transcription with turn detection. Deepgram says it can cut agent response latency by 200–600 ms compared with pipeline approaches, with a p90 of 1 second and a p95 of 1.5 seconds. Those are vendor numbers. Flux accepts 8 kHz mu-law audio directly and extends to 10 languages in its multilingual variant. The catch is that Flux redacts numbers only, replacing each span with a single asterisk. A voice agent can live with that. A compliance recording cannot.
AssemblyAI's Universal-3.5 Pro Realtime claims end-of-turn detection around 300 ms, with P50 latency of 546 ms in balanced mode. That is also vendor-run. One billing detail matters on a contact-centre floor. AssemblyAI bills streaming on how long the WebSocket stays open, not on the audio sent, so an agent-assist session left open through a hold is billed for the hold. OpenAI's gpt-live-transcribe is $0.017/min, more than double the others. Its realtime transcription sessions are also not eligible for Zero Data Retention, even though the file-transcription endpoint is.
Capacity is the quieter differentiator. Deepgram publishes up to 150 concurrent streaming connections on pay-as-you-go and 225 on Growth. AssemblyAI publishes a rate on new sessions rather than a concurrency ceiling: 100 new streams per minute on pay-as-you-go. A 300-seat floor streaming both channels needs 600 concurrent streams, four times Deepgram's pay-as-you-go ceiling, so ask each vendor for the concurrency in writing before you sign.
Stereo recordings double the API bill, and diarisation is the GPU cost
Transcribing both channels of a stereo call gives perfect speaker attribution, but APIs bill it twice. Mixing down to mono and diarising halves the API bill but makes speaker labels worse. For a self-hoster, the maths runs the other way.
AssemblyAI is explicit: when multichannel transcription is enabled, "each channel is transcribed and billed separately". For Deepgram, a user testing multichannel in Deepgram's own forum reported that billing multiplies by the number of channels, even for channels filled with silence. No Deepgram staff member has answered that thread. Ask for the answer in writing.
The alternative is to mix down to mono and pay for diarisation, meaning software that works out who spoke when. Its quality on phone audio is the weak link. Pyannote's own model card puts the diarisation error rate on CALLHOME, a telephone-conversation benchmark, at 26.7% for its free community-1 model against 16.6% for its commercial precision-2. AssemblyAI reports its best telephony cpWER (a word error rate that also penalises wrong speaker labels) at 17.18 to 17.78, and says it beats Deepgram Nova-3 on average, 30.17 against 37.92. Both figures are from AssemblyAI's own benchmark. Diarisation is included free on Deepgram's pre-recorded tier and costs +$0.02/hr on AssemblyAI batch.
The self-hosting surprise is that diarisation, not transcription, eats the GPU. Pyannote states a real-time factor of about 2.5% on a V100, roughly 1.5 minutes of processing per hour of audio. Parakeet's long-form throughput of 1,000× real time on an A100 works out to about 3.6 seconds per hour. The two figures come from different hardware, a V100 plus a CPU for clustering against a batched A100, but even allowing for that, diarising a call takes well over ten times the compute of transcribing it with Parakeet. Pyannote's benchmark figure is also optimistic. One user on a cloud V100 measured 12 minutes to diarise a 10-minute file, with the CPU pinned at 100% and the GPU idle.
So if you self-host, keep the stereo split and skip diarisation entirely. If you buy an API, price both configurations, because mono plus diarisation is roughly half the bill.
Where does self-hosting speech-to-text actually pay?
With a two-engineer team, self-hosting Parakeet breaks even at roughly 1.8 to 2.4 million call minutes a month against an API transcribing both channels, and 3.6 to 5.5 million against an API's cheaper mono-plus-diarisation option. Self-host Whisper instead and each line moves further out.
First, the GPU. Modal lists an A100 80GB at $0.000694 per second, which is $2.50 an hour. At the Open ASR long-form throughput, that is about $0.0025 of GPU per audio-hour on Parakeet and $0.036 on Whisper. Your 5 million stereo minutes are 166,700 channel-hours a month, so the monthly GPU bill comes to about $420 on Parakeet and about $6,070 on Whisper at full load. Parakeet's compute rounds to zero. Whisper's does not, and it is sensitive to utilisation: run the same fleet at 25% and Whisper's GPU line alone reaches about $24,300 a month.
Second, the people. The BLS median wage for software developers was $135,980 in May 2025. Two engineers at that figure cost $22,660 a month in wages alone, before benefits. That is a floor. ML infrastructure engineers earn above the median, and two is the minimum, because one person cannot carry an on-call pager for a production service indefinitely.
Third, the API bill for the same 5 million call minutes, at list, with PII text redaction:
| API configuration | Per call-minute | Monthly at 5M min | Parakeet break-even | Whisper break-even |
|---|---|---|---|---|
| Deepgram Nova-3 pay-as-you-go, stereo | $0.0126 | $63,000 | 1.8M min | 2.0M min |
| AssemblyAI Universal-3.5 Pro, stereo | $0.0097 | $48,300 | 2.4M min | 2.7M min |
| Deepgram Nova-3 pay-as-you-go, mono + diarisation | $0.0063 | $31,500 | 3.6M min | 4.5M min |
| Deepgram Growth, mono + diarisation | $0.0053 | $26,500 | 4.3M min | 5.5M min |
| AssemblyAI Universal-3.5 Pro, mono + diarisation + redaction | $0.0052 | $25,800 | 4.5M min | 5.7M min |
| AssemblyAI Universal-2, mono + diarisation + redaction | $0.0042 | $20,800 | 5.5M min | 7.7M min |
The inputs come from the pricing pages. Deepgram: $0.0043/min pre-recorded, $0.0036 on Growth, redaction +$0.0020 (+$0.0017 on Growth), diarisation included. AssemblyAI: $0.21/hr Universal-3.5 Pro, $0.15/hr Universal-2, diarisation +$0.02/hr, PII text redaction +$0.08/hr. The stereo rows assume per-channel billing, including add-ons. The self-host side is two engineers plus full-load GPU on stereo audio with no diarisation. The API "mono" rows include diarisation, which the self-hosted stereo pipeline does not need.
Three things move these lines:
- Negotiated rates move them right. Deepgram's Growth plan is a $4K+ annual commitment offering up to 20% savings, Enterprise is contact-sales, and AssemblyAI advertises volume discounts. At 5 million minutes you will not pay list.
- Reserved GPUs move them right by about 15%. Two A100s held around the clock for redundancy add roughly $3,650 a month over per-second serverless.
- Deepgram's training opt-out may move Deepgram's rows left, by a lot. See the next section.
There is also a third path, where you license the vendor's model and run it in your own VPC. Self-hosted Deepgram requires an Enterprise plan and a licence server that reports usage. AssemblyAI's self-hosted Universal-3.5 Pro uses "the same usage-based pricing as our cloud service". That makes it a residency purchase, not a saving. It is the same conclusion this site reached on self-hosting a vector database for residency rather than the bill.
Who keeps your call audio, and who trains on it?
Deepgram trains on your audio by default under a perpetual licence. AssemblyAI trains on some redacted files by default but lets paid customers opt out for free. OpenAI does not train on API data by default. For a contact centre, this is the column most likely to end the evaluation.
Deepgram. Its terms, last updated 6 August 2026, let it use submissions "for other business purposes, including training and testing our Models." They grant Deepgram "an irrevocable, perpetual, transferable, sublicensable (through multiple tiers), fully paid, royalty-free, nonexclusive and worldwide right" to your content, which survives termination. You opt out per request with mip_opt_out=true, after which "data from opted-out requests is retained only for the duration necessary to process the request". The price of opting out is the problem. A Deepgram collaborator said in June 2025 that opting out "forgoes a 50% discount". The pricing page checked today does not state an opt-out rate. If that statement still holds, the list rate in the tables above is the opted-in rate, and the opted-out rate is double. Get the opted-out rate in writing before you compare Deepgram to anything.
AssemblyAI. By default, uploaded audio is deleted between 24 and 48 hours, final transcripts from 30 days, and "only certain files … are used for model training", after a PII redaction pass. The opt-out, TTL and BAA are all "self-serve at no additional cost" on paid plans. Free accounts cannot opt out, so do not run a pilot on real calls from a free account.
OpenAI. API data is not used for training unless you opt in, abuse-monitoring logs are kept up to 30 days, and /v1/audio/transcriptions is eligible for Zero Data Retention. For a governance team that has already approved OpenAI, GPT-Transcribe is the path of least resistance. Its gaps are functional ones: no redaction, and diarisation only through a separate model.
Self-hosted. The audio never leaves your network, which is the one argument no API can match at any price. If a regulator asks which jurisdiction ran the inference, the residency questions from our LLM buyer's guide apply here unchanged.
A transcript is a queryable copy of the card number
Transcribing a payment call converts audio that is hard to search into text that is easy to search, and that brings it into the part of PCI DSS your recording vendor was careful to stay out of. Redaction quality is therefore a compliance control, not a feature.
The PCI Security Standards Council's guidance on telephone payments says it is prohibited to store CAV2, CVC2, CVV2 or CID codes after authorisation in digital audio "if that data can be queried". A transcript can be queried by definition. If your ASR mishears one digit of a spoken security code, the redaction layer that relies on the transcript may miss it.
The vendors differ here in ways a feature list hides. Deepgram's entity redaction is "applied for English only"; for other languages it redacts numbers only. It also warns that no_delay=true trades "redaction performance" for latency. AssemblyAI offers 60+ policies, including credit_card_cvv, across 50 languages, with beeped or silenced audio redaction. It also warns that redaction only covers the text field, so summaries and entity-detection output "may still include PII." If you self-host, the default open tool is Microsoft's Presidio, whose README states plainly that there is "no guarantee that Presidio will find all sensitive information". Whichever route you take, test redaction on 50 recorded payment calls before anything else, because this is the failure that ends up in front of your acquiring bank.
Who should not pick each of these
The part a vendor page will never tell you.
- Do not pick AssemblyAI if you need self-serve streaming concurrency published up front, or if your agent-assist sessions idle through long holds. Session-duration billing charges for the silence. Its best accuracy numbers belong to a model it has since replaced, so re-run them yourself on Universal-3.5 Pro.
- Do not pick Deepgram if your legal team will not accept a perpetual, sublicensable licence to customer audio and your finance team will not pay the opt-out premium. Skip it too if you need entity redaction in Spanish, Hindi or Tagalog. Also skip it for compliance recording on Flux, which has asterisk-only redaction.
- Do not pick OpenAI GPT-Transcribe if you need redaction or speaker labels in the same call. OpenAI's own migration guide tells you to keep
whisper-1if you depend on native SRT or VTT output. - Do not pick NVIDIA Parakeet TDT 0.6B v3 if your callers speak anything outside its 25 European languages, if you are under the crossover, or if nobody on your team has run a GPU service in production. Pull the weights from Hugging Face or run them on Modal, but the pager stays with you.
- Do not pick OpenAI Whisper large-v3 for English contact-centre audio. The one case where it still earns a place is language breadth: 99 languages against Parakeet's 25. Run it behind VAD and never on raw split-channel audio.
The loser is self-hosted Whisper. The model is good, but it is the most expensive way to get the least accurate result on your specific audio. It is 15.0% against 9.2% on accented service calls, needs about 14 times the GPU of Parakeet on long recordings, invents sentences when the line goes quiet, and has 3 seconds of streaming lag. Teams pick it because it is the open model they have heard of, and that is not a reason.
The decision, in the order that predicts regret
This week. Pull 200 real calls, stratified by accent, line quality and hold time, and have someone hand-correct the transcripts. Score WER, and separately score entity recall on account numbers and names. Every number on this page is subordinate to that set.
This week. Find out whether your platform records stereo. If it does, and Amazon Connect does, every API quote you receive should be priced both ways: per channel, and mono plus diarisation.
This month. Get three numbers in writing from Deepgram: the opted-out rate, how multichannel is billed, and your concurrent-stream ceiling. The whole comparison depends on the first one.
This month. Run the 50-payment-call redaction test on every candidate. Count missed CVVs, not average F1.
Before the contract. If you are past about 3 million stereo minutes a month and have two engineers who have run GPU services before, stand up Parakeet on the same 200 calls and price it honestly, pager included. The GPU scheduling mistakes that strand capacity apply to ASR fleets too.
What changes the answer: a live voice agent where turn-taking decides containment (favours Deepgram Flux); a residency clause (self-host, or licensed self-host from either vendor); non-European languages (Whisper back in for self-hosting, or an API); or an independent 8 kHz, multi-accent benchmark that includes the commercial APIs, which nobody has published. The day someone does, re-run everything.
The Bottom Line
Speech-to-text got cheap enough that the per-minute rate has stopped being the decision. At a half-cent a minute, the real costs sit in things that never appear on the rate card: a channel count that silently doubles the bill, a diariser that costs over ten times the compute of the transcriber, a training licence that outlives the contract, and a redaction layer that is only as good as the digit the model heard. The same thing happened with cloud storage a decade ago. Teams negotiated hard on the price per gigabyte and discovered the real bill was egress and API calls.
Buy AssemblyAI for the recordings, buy Deepgram only where a voice agent needs Flux, and self-host only when your volume and your team can carry it. When you do self-host, pick the model that was benchmarked on your kind of calls, which is not the one with the most GitHub stars.
Whisper is the model everyone has heard of. Parakeet is the one that heard your callers.
Continue Reading
Best AI Contact Center Platforms: Containment Isn't Resolution Eleven Voice Agents, One Bank Call, No Clean Winner Voice AI's Rent vs. Own Moment: Mistral Bets on Open Weights Alianza Bought Skribby. Your Meeting Audio Changed Owners. LLM Data Residency: Which Providers Actually Keep Data In Region OpenAI vs Cohere vs Qwen3: Re-Embedding 10M Chunks Costs $520
