Inception's Mercury Voice is worth a bake-off slot in your voice-agent pipeline, not a model swap. The diffusion LLM, generally available to enterprise customers since September 29, lists at $0.40 per million input tokens and $1.50 per million output, is 50% off at launch, and works out to "about $0.009 per minute of conversation" by Inception's own math. Its headline latency — 320ms median, 750ms p95 — was measured on the model alone, on Inception's own test set, and the benchmark win it claims arrives without a single number attached in the launch text. The price is real. The quality claim is yours to prove.
That is the whole decision. The rest of this piece is how to make it without getting burned.
What Did Inception Actually Ship?
Mercury Voice is a diffusion language model tuned for the LLM slot of a voice agent — the part between speech-to-text and text-to-speech. A diffusion LLM (dLLM) generates by refining many tokens in parallel rather than emitting one token at a time, which is where Inception's speed story comes from. Per the launch post, the model has a 128K-token context, up to 50K output tokens, three reasoning-effort settings (low, medium, high), and it "drops into" LiveKit, Pipecat, Vapi, Retell or a custom stack through an OpenAI-compatible endpoint.
Access is gated. Inception's model documentation lists the ID as mercury-voice, marks it "Enterprise access only," and routes buyers to sales. The same page shows a cached-input rate of $0.02 per million tokens at the launch discount ($0.04 list) — a detail that matters more than the headline price for voice agents, which resend a long system prompt on every turn.
Three customers are named. Audivi AI runs drive-through ordering, Altur negotiates payment plans on calls for financial institutions, and OpenCall reports a median model response latency near 170 milliseconds. Inception is not a newcomer to enterprise money: the same report notes a $50 million seed round led by Menlo Ventures, with M12 and Snowflake Ventures participating.
How Fast Is 320ms, Really?
320ms is time to first answer token from the model, not the delay your caller hears. Inception defines it as "how long until the model starts producing the words," measured "on a test set of real customer-service prompts" — its own set, with no hardware, region or prompt-length disclosure in the post. AlphaSignal's write-up adds the condition that matters: the 320ms figure is at low reasoning effort, and time to first answer token "does not capture the caller's complete mouth-to-ear delay."
Now put it in the budget. Inception frames the target as responding "within about 500 ms of the caller finishing". Practitioners who build these pipelines set the whole round trip at roughly 800 milliseconds before a conversation starts to feel slow, and they name speech-to-text and LLM inference as the two stages that "consistently account for most of the round trip." Turn detection, TTS time-to-first-audio, WebRTC or PSTN transport and the jitter buffer all sit on top of the model's 320ms.
So the median is good news. The p95 of 750ms is the number to stare at: on one call in twenty, the model alone consumes almost the entire end-to-end budget before a single phoneme is synthesised. Contact-center callers remember the one long pause, not the median.
The comparisons are also carefully chosen. Inception says Mercury Voice is "5.9x faster than GPT-6 Luna (no reasoning)" — a comparison against a competitor with reasoning off, while its own headline figure is at low effort. Neither is wrong. Neither is your workload.
Where Is the Benchmark Score?
The launch post names five public benchmarks and, in its text, reports no score on any of them. The claim is that Mercury Voice "outperforms newer models, such as GPT-6 Luna, GLM-5.3-Flash, and Gemini 3.5 Flash-Lite, on a composite benchmark that includes τ³-bench Telecom, τ³-bench Retail, τ³-bench Airline, IFBench, and BFCL v4", and that it "achieves higher quality than almost every other model we tested." A composite, averaged, against a hand-picked field, with "almost" doing quiet work. AlphaSignal flags all of it as vendor-reported, without independent verification.
The steel-man: the benchmarks are public, so Inception is inviting you to check. That is true, and it is why you should. But each has a known weakness a buyer should price in.
- BFCL v4 is noisy. Epoch AI sampled 50 of its 5,088 scored items on September 10, 2026 and found defects that could affect accuracy in 24 of them — 48%, with the memory category worst at 70%. A small composite win built partly on BFCL can be noise.
- τ³-bench leaderboards mix provenance. One aggregator lists Mercury 2.5, Inception's general model, at 96.0% on τ³-bench, above Claude Opus 4.5 at 70.2% — and the same page warns that provider tables without matched controls "are useful source receipts, not one apples-to-apples leaderboard." A 26-point lead over a frontier model is a reason to re-run, not to believe.
- Text scores do not survive the microphone. Sierra's own τ-Voice paper found that where GPT-5 with reasoning reached 85% on text tasks, voice agents reached 31–51% in clean audio and 26–38% with noise and accents, and that 79–90% of failures stemmed from agent behaviour. Sierra's follow-up shows the best voice result climbing from 30% in August 2025 to 67% in April 2026 — progress, but still well short of text.
Mercury Voice is a text model sitting between ASR and TTS. Its τ³-bench text score, whatever it is, tells you how it reasons over a clean transcript. Your callers do not produce clean transcripts.
Does $0.009 a Minute Change Your Unit Economics?
It changes the LLM line, which is a minority of what a voice minute costs. Inception prices Mercury Voice at "~5x cheaper than GPT-4.1 ($0.045 per min)", but the post does not disclose the tokens-per-minute or input/output ratio behind "a typical voice-agent profile" — so re-derive it from your own call logs before it goes in a business case. The launch discount carries no stated end date.
Now the rest of the minute. Vapi's published rates put its hosting fee at $0.05 a minute, Deepgram transcription at $0.0095–$0.0099, ElevenLabs voice at $0.0146–$0.0238, and Twilio outbound telephony at $0.014, with OpenAI models passed through at $0.0077–$0.0452. Taking the low end of each non-LLM line, a Vapi minute costs about $0.088 before the model (my arithmetic, from those listed rates). Add GPT-4.1 at Inception's $0.045 and the minute is roughly $0.133; swap in Mercury Voice at $0.009 and it is roughly $0.097. That is a ~27% cut in per-minute cost — about $36,000 a month at a million minutes. Material. Not transformative, and entirely erased if containment drops a single point and those calls roll to a human.
There is a second pricing wrinkle. Inception's Mercury 2.5 announcement lists its general model at $0.20/$0.75 per million tokens with an 80% launch discount to $0.04/$0.15 — and credits the same OpenCall deployment with a median model response latency "close to 170 milliseconds." Mercury Voice's discounted price is Mercury 2.5's list price. Before you pay the voice premium, ask Inception which model produced OpenCall's 170ms, and put Mercury 2.5 in the bake-off too.
What to Do About It
Treat Mercury Voice as a candidate in a controlled bake-off, run on your transcripts, measured end to end. The model slot in LiveKit, Pipecat, Vapi and Retell is designed to be swappable, which makes this cheap to test and expensive to skip.
This Week:
- Pull 200 real call transcripts from your worst-performing intent — the authentication, payment or cancellation flows where agents escalate most — and freeze them as a replay set with the tool schemas they actually call.
- Email Inception's sales team for three things in writing: the per-benchmark scores behind the composite, the token profile behind $0.009 a minute, and the launch-discount end date. A vendor confident in its numbers will send them.
This Month:
- Run the replay through your existing pipeline twice — incumbent model versus Mercury Voice (and Mercury 2.5) — at matched reasoning effort. Log mouth-to-ear p50 and p95, tool-call accuracy, and task completion. Size the set properly first; small evals produce gaps that are not real.
- Re-run τ³-bench Telecom and Retail yourself on the candidates, and drop any BFCL v4 category Epoch flagged before comparing.
Before Renewal:
- Price the minute, not the token. Rebuild your per-minute cost with your STT, TTS, telephony and platform fees, then weight it by containment. A cheaper model that escalates one more call in a hundred is a more expensive model.
The Bottom Line
Mercury Voice moves the cheapest line on your voice bill and makes a claim about the most expensive one — whether the call gets resolved — that nobody outside Inception can yet check. This is the same pattern the speech-to-text market went through: vendor word-error-rate charts first, then buyers learning that their own accents, noise and jargon decided the winner. The eleven-stack bank-call test produced no clean winner for the same reason.
A fast, cheap model in the LLM slot is genuinely useful. Speed is the easy part to measure. Resolution is the part you are paying for.
The latency is a spec sheet. The benchmark win is a sentence. Your call logs are the only benchmark that renews.
Continue Reading
- Eleven Voice Agents, One Bank Call, No Clean Winner
- Best AI Contact Center Platforms: Containment Isn't Resolution
- Qwen-Audio 3.1 Transcribes an Hour of Calls for About Two Cents
- GPT-6 Sol Halves the Token Price but Benchmarks It at Top Effort
- LiveKit Buys Loophole Labs, Leaving Architect Buyers Without a Word
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
