Microsoft's MAI-Transcribe-2-Streaming Tops WER With No SLA

Microsoft's MAI-Transcribe-2-Streaming claims a 2.5% streaming WER at $0.54 an audio hour, but it ships as a public preview with no SLA and no word-level timestamps, at 2.7 times Grok's streaming price.

By Rajesh Beri·October 3, 2026·9 min read
Share:
A contact-center headset resting on a desk next to a laptop showing a live scrolling transcript, with a paper contract page beside it marked 'preview' in red pen, under office lighting.

Illustration generated using AI

Microsoft's streaming transcription model posts the best accuracy in the category, but you cannot put it under a production SLA yet, and on most contact-center workloads it costs more than the models it beats. MAI-Transcribe-2-Streaming, released October 1, claims a 2.5% word error rate with final text arriving 0.13 seconds after the speaker stops, at $0.54 per audio hour through the end of 2026. Microsoft's own documentation says it is a public preview "provided without a service-level agreement" and "not recommended for production workloads." If you run real-time speech-to-text for a contact center or a voice agent, test it against Gemini 3.5 Transcribe Live and Grok Voice Transcribe 2.0 on your own call audio, and write your renewal plan around the price after December 31.

Real-time (streaming) speech-to-text is a model that transcribes audio while the speaker is still talking, emitting provisional "partial" text and then a confirmed "final" segment. Agent-assist screens, live captions and voice agents all depend on it, and its latency is part of every voice agent's response time.

What Microsoft Actually Shipped

Microsoft shipped its first streaming transcription model, plus two text-to-speech models, and priced the transcriber at a promotional rate. The Microsoft AI announcement says the model ranks first on Artificial Analysis for accuracy on both final and partial transcripts, produces first text in "just over 100ms," supports 60 languages with continuous automatic language detection, and costs $0.54 per audio hour as an introductory rate through the end of the year. The same post launched MAI-Voice-2.1 at $22 per million characters and MAI-Voice-2.1-Flash at $15 per million.

The benchmark detail comes from the trade write-ups. MarkTechPost reports a 2.5% WER on both final and first-partial transcripts, 0.13 seconds to final and 0.12 seconds to first partial, and a "#1 of 38 models" ranking attributed to Artificial Analysis. That matching partial-and-final accuracy matters for agent-assist, where the agent reads the partial text before the model confirms it.

Artificial Analysis has since posted the result under its own name, putting the model first on its AA-WER Streaming index ahead of Grok Voice Transcribe 2.0 at 2.7%. A benchmark set is still not your call audio, and OrcaRouter's release analysis points out that Microsoft has published no streaming diarization error rates, no per-language breakdown and no far-field noise results.


Is MAI-Transcribe-2-Streaming Production Ready?

No, not by Microsoft's own terms. Both the model overview and the Speech SDK guide, updated October 1, carry the standard preview notice: no service-level agreement, not recommended for production, and "certain features might not be supported or might have constrained capabilities." Your contract follows the Microsoft Learn label.

The preview terms carry more than an uptime gap. The Supplemental Terms of Use for Microsoft Azure Previews say Microsoft may process and store customer data submitted to AI services to monitor for abusive or harmful uses. For recorded customer calls carrying card numbers or health details, that sentence goes to your privacy and compliance leads before any pilot.

Three engineering limits show up in the SDK guide:

  1. Results "don't include detected language information, confidence scores, or word-level timestamps," and only final results carry segment-level offset and duration. If your PCI redaction mutes card numbers in the recording by word timing, or your QA tooling flags low-confidence words for review, this model cannot feed either today.
  2. The service disregards the SDK's ProfanityOption and OutputFormat, so any masking has to move into your own pipeline.
  3. Microsoft says you can access the model globally and that Azure serves it from Sweden Central, Central US and Southeast Asia, with East US 2 "coming soon." A US-only data residency commitment needs a written answer on where your audio is processed.

How It Compares on Price and Accuracy

On accuracy it leads the published field by a small margin; on price it is the most expensive of the serious contenders except Gemini, which costs about the same. The table uses each vendor's list price for streaming and the WER figures cited in the release coverage.

Model Streaming WER Latency to final Streaming price per audio hour Status
MAI-Transcribe-2-Streaming 2.5% 0.13s $0.54 (intro, to Dec 31) Public preview, no SLA
Grok Voice Transcribe 2.0 2.73% 0.49s $0.20 Released Sep 18
Gemini 3.5 Transcribe Live 4.0% 0.40s about $0.54 (token-billed) No preview label
ElevenLabs Scribe v2 Realtime 3.59% 0.14s not compared here n/a
Deepgram Flux (English) not compared here n/a $0.39 (promo) n/a

Sources: OrcaRouter for Grok, ElevenLabs and Gemini WER and latency; DataNorth for Grok pricing; Google's pricing page for Gemini; Deepgram's pricing page for Flux.

Google bills Gemini 3.5 Transcribe Live at $3.50 per million audio input tokens and $21 per million text output tokens, which Google estimates blends to about $0.009 a minute, or roughly $0.54 an hour. Google's pricing page does not label it a preview. Its launch post gives 4.0% streaming WER, 85+ languages and custom vocabulary biasing of up to 1,000 terms. OrcaRouter's head-to-head flags a 10-minute session cap on Google's Live API, which a long support call will hit.

xAI prices Grok Voice Transcribe 2.0 at $0.20 an hour streaming and $0.10 batch, the same as version 1.0. DataNorth lists word-level timestamps, multichannel transcription up to eight channels and keyterm prompting, with limits of 100 concurrent streaming sessions and service from US-East-1. It also notes xAI published no 1.0-versus-2.0 accuracy table of its own.

Deepgram lists Flux English at $0.0065 a minute pay-as-you-go (a promotion off $0.0077) and Nova-3 monolingual at a promotional $0.0048 (regular $0.0077), with streaming diarization an extra $0.0020 a minute.


What the Accuracy Gap Is Worth

The accuracy lead over Grok is about two errors per thousand words, and you pay 2.7 times Grok's list price for it. At 2.5% WER a 10,000-word transcript carries about 250 errors; at Grok's 2.73% it carries about 273; at Gemini's 4.0% it carries about 400. Against Gemini, Microsoft cuts errors by more than a third for roughly the same money. Against Grok, the gain is small enough that your own audio, with its accents, line noise and product names, will decide it.

The bill tells a different story. One million streamed minutes a month is about 16,667 audio hours. At list prices:

  • MAI-Transcribe-2-Streaming: about $9,000 a month at the introductory rate.
  • Gemini 3.5 Transcribe Live: about $9,000, on Google's blended estimate.
  • Deepgram Flux English: about $6,500 at the promotional rate.
  • Deepgram Nova-3 with streaming diarization: about $4,800 plus $2,000, so $6,800.
  • Grok Voice Transcribe 2.0: about $3,333.

Microsoft's own batch model is the other comparison. Batch MAI-Transcribe-2 costs $0.10 an hour, so streaming costs 5.4 times as much. If a workload does not need text while the call is live (post-call QA, summaries, compliance search), sending it through streaming wastes most of the spend.

The price you are quoted today also expires. Microsoft calls $0.54 an introductory rate through the end of 2026, and the Azure Speech pricing page labels MAI-Transcribe-2 a "limited-time promotional offer till 12/31/2026" that "reflects public preview pricing." Microsoft has not said what the rate becomes. We saw the same pattern with Gemini 4 Argon's introductory pricing and missing GA date: a leaderboard number arrives first and the commercial terms arrive later.

Who Should Test It Now

Voice-agent teams where latency is the binding constraint have the strongest case. Final text at 0.13 seconds, against 0.40 for Gemini and 0.49 for Grok, takes 0.27 to 0.36 seconds off every turn, and our voice AI per-minute pricing teardown shows how quickly the speech layer stacks up inside a platform's per-minute rate. Teams already on Microsoft Foundry can try it with Speech SDK v1.52.0 or the OpenAI Realtime-compatible WebSocket API, so the integration cost of a test is low.

Agent-assist and live-captioning teams come next, because matched partial and final accuracy means fewer corrections flicker on screen.

Post-call analytics teams should not move. Batch transcription is a fifth of the price, and our Qwen-Audio 3.1 piece found batch prices at about two cents an hour. Teams that depend on word timestamps for redaction, or on diarization for QA scoring, should wait until Microsoft documents both. Our Deepgram vs AssemblyAI vs Whisper comparison covers why those two features usually decide a contact-center STT choice.


What to Do About It

This Week:

  1. Pull 200 recorded calls that reflect your real mix (accents, hold music, speakerphone, your product names) and get them transcribed by a human, or reuse a gold set you already have.
  2. Ask your Microsoft account team three questions in writing: when MAI-Transcribe-2-Streaming leaves preview, what the per-hour rate becomes on January 1, and which region processes US audio.
  3. Send the Azure preview data-handling clause to your privacy lead before any customer audio touches the endpoint.

This Month:

  1. Run the 200 calls through MAI-Transcribe-2-Streaming, Gemini 3.5 Transcribe Live, Grok Voice Transcribe 2.0 and your incumbent (Deepgram, AssemblyAI or Azure Speech). Score WER on your product terms separately from the overall rate.
  2. Measure latency from your own region, not the vendor's figure, and note where any provider drops a call longer than its session cap.
  3. Split your traffic: anything that does not need live text goes to batch at $0.10 to $0.20 an hour.

Before Renewal:

  1. Do not sign a 2027 streaming commitment priced on a promotional rate. Get the post-December price in the order form or a cap on its increase.
  2. Make GA status, an uptime SLA and word-level timestamps contract conditions for any move of production traffic.
  3. Keep a second provider integrated. Our model deprecation notice-window analysis shows how much warning vendors give before a model changes underneath you.

The Bottom Line

Plan for a familiar sequence: a new model posts the best WER first, and the features contact centers rely on (timestamps, diarization, redaction hooks, an SLA) follow later, if at all. Microsoft now has the best published streaming accuracy and the lowest latency in the category, priced level with Google and above xAI and Deepgram, with no SLA and no list price for January. Our contact center platform guide measures vendors on verified resolution, and transcription accuracy feeds that number. Run your own 200 calls through all four models before you move any production traffic.

Continue Reading

Share:

Frequently Asked Questions

How much does MAI-Transcribe-2-Streaming cost?

Microsoft prices it at $0.54 per audio hour, about $9 per 1,000 minutes, as an introductory rate through the end of 2026. Microsoft has not published the rate after December 31. The batch MAI-Transcribe-2 model costs $0.10 an hour.

Is MAI-Transcribe-2-Streaming generally available?

No. Microsoft Learn lists it as a public preview provided without a service-level agreement and not recommended for production workloads.

How does MAI-Transcribe-2-Streaming compare with Gemini 3.5 Transcribe Live and Grok Voice Transcribe 2.0?

Artificial Analysis ranks it first on streaming accuracy at 2.5% WER, with final text 0.13 seconds after speech ends. Release coverage puts Grok Voice Transcribe 2.0 at 2.73% and 0.49 seconds for $0.20 an hour, and Gemini 3.5 Transcribe Live at 4.0% and 0.40 seconds for about $0.54 an hour.

Does MAI-Transcribe-2-Streaming return word-level timestamps?

Not today. Microsoft's Speech SDK guide says results do not include word-level timestamps, confidence scores or detected language, and only final results carry segment-level offset and duration. That limits audio redaction and QA tooling that depend on word timing.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →