Qwen-Audio 3.1 Transcribes an Hour of Calls for About Two Cents

Qwen-Audio 3.1 bills speech recognition in tokens: at 25 tokens per audio second it works out to about $0.02 an hour in Singapore, against $0.15-$0.26 at AssemblyAI and Deepgram. The catches are a five-minute request cap, per-channel billing and a residency decision.

By Rajesh Beri·September 27, 2026·10 min read
Share:
A contact-centre headset resting on a desk beside a printed invoice and a calculator, a stopwatch lying on top of the invoice, under office lighting.

Illustration generated using AI

Alibaba's new speech-recognition model transcribes an hour of audio for roughly two cents through its international region — about one-eighth of what AssemblyAI lists and one-thirteenth of Deepgram's pay-as-you-go rate. You will not find that number on the pricing page, because Qwen-Audio 3.1 bills in tokens, not minutes. The conversion is simple once you know it, the savings are real, and the catches — a five-minute request cap, no batch tier, two pricing regions, and no diarization in the base model — are exactly the ones a contact-center buyer needs to price in before using this as leverage at a speech-vendor renewal.

Alibaba shipped five models under the Qwen-Audio 3.1 name on September 23 — ASR, ASR-Next, TTS, TTS-Next and a realtime model — and cut prices by up to 95% for ASR, about 70% for TTS and about 85% for realtime. "Up to" is doing work in that sentence. Here is what the meter actually says.

What Does Qwen-Audio 3.1 ASR Cost per Hour?

About $0.019 per audio hour in Singapore and about $0.015 in Beijing, on a typical conversational transcript. The arithmetic has three inputs.

First, the rate card. The qwen-audio-3.1-asr-flash model page lists $0.15 per million input tokens and $0.47 per million output tokens in Singapore, the international region, and $0.113 and $0.382 in Beijing.

Second, the conversion. Alibaba's Qwen-ASR API reference states that "each second of audio is converted to 25 tokens," with anything under a second rounded up to one. An hour of audio is therefore 90,000 input tokens. One caution: that reference documents the Qwen3-ASR-Flash family, and the 3.1 model page does not restate the rate. The 3.1 page's 7,168-token input ceiling works out to 287 seconds at 25 tokens per second — consistent with the five-minute limit Alibaba documents for its short-form ASR — but confirm the rate on your own first invoice. The API returns both audio_tokens and seconds in its usage block, so it takes one call to check.

Third, the output. A transcript is text, and text is billed at the output rate. Conversational speech runs at roughly 150 words a minute; call it 12,000 output tokens per hour. That figure is our estimate, not Alibaba's.

Per audio hour Input Output (est.) Total
Qwen-Audio 3.1 ASR, Singapore 90k × $0.15/M = $0.0135 12k × $0.47/M = $0.0056 ≈ $0.019
Qwen-Audio 3.1 ASR, Beijing 90k × $0.113/M = $0.0102 12k × $0.382/M = $0.0046 ≈ $0.015
AssemblyAI Universal-2 (async) — — $0.15
AssemblyAI Universal-3.5 Pro (async) — — $0.21
Deepgram Nova-3 pre-recorded, PAYG $0.0043/min — $0.258
Deepgram Nova-3 pre-recorded, Growth $0.0036/min — $0.216

Input dominates, which matters: a speaker who talks faster moves the total by fractions of a cent, not multiples. Prices are as listed on each vendor's page on September 27, 2026.

For scale, a contact center transcribing 100,000 call hours a month would pay about $25,800 at Deepgram's pay-as-you-go rate, $15,000 at AssemblyAI Universal-2, and under $2,000 on Qwen-Audio 3.1 in Singapore — before any add-ons on either side.


Is "Up to 95%" the Number You Should Use?

No — measured against Alibaba's own previous international price, the Singapore cut is roughly 84%. The prior Qwen3-ASR-Flash model is billed by duration at $0.000032 per second in both Singapore and Beijing, per its model page as of September 27, 2026. That is about $0.115 an hour. Against $0.019, the reduction is about 84%. (A pricing correction in Portkey's model registry recorded $0.000035 per second, which would put the cut nearer 85%.)

That is still a very large cut. But "up to 95%" is a headline number, and the reader who walks into a renewal quoting it will be corrected by a vendor sales engineer who has done this math. Quote the per-hour figure you computed, and show the working.

What Does Token Billing Hide From a Contact-Center Buyer?

Three things: the request ceiling, the channel count, and the output cap. None is disqualifying. All three change the architecture you would build.

The five-minute ceiling. The 3.1 flash model accepts at most 7,168 input tokens per request, per its model page — under five minutes of audio. A seven-minute support call has to be chunked, and chunk boundaries are where words get split and punctuation breaks. Alibaba's speech recognition guide describes a separate Qwen-Audio-3.1-ASR-Flash-Filetrans variant that takes recordings up to 12 hours and 2 GB, which is the right tool for call archives — but the model page we read prices only the flash model. Get the filetrans rate in writing before you model it.

Channels are billed separately. Contact-center audio is often recorded in stereo, agent on one channel and customer on the other, which is a cheap way to get speaker separation. The API reference, in its section on the Filetrans model's channel_id parameter, notes that "each specified audio track is billed separately"; expect the same rule wherever you submit both channels. Two tracks, two sets of input tokens: the Singapore hour goes to roughly $0.03. Still an order of magnitude below the incumbents, but budget it.

Diarization is a different model. Speaker identification with timestamps belongs to ASR-Next, not the base ASR model, according to The Decoder's rundown, and ASR-Next was listed as "coming soon" at launch. For comparison, AssemblyAI charges an extra $0.02 an hour for async diarization. If your QA workflow depends on mono recordings plus diarization, you cannot price Qwen for that workload yet.

A smaller point: the output cap is 1,024 tokens per request. At our 12,000-tokens-per-hour estimate, a full 287-second chunk produces about 960. A fast talker could, in principle, press against that ceiling. Test it with your fastest agents' recordings.

What the Model Page Says It Won't Do

The 3.1 flash model supports neither batch inference, fine-tuning nor context caching, according to the "unsupported features" line on its model page. The rate limit is 600 requests per minute in both regions.

No batch tier means there is no further discount for overnight archive jobs — the price you see is the floor. No fine-tuning means your product names, drug names and account codes live or die on the base model's vocabulary; the model page advertises "context enhancement," which is the lever to test instead. And 600 requests per minute of sub-five-minute chunks caps a single workspace at roughly 48 hours of audio per minute of wall-clock time, which is plenty for most, and worth checking for anyone doing a one-time backfill of years of recordings.

On accuracy, the only figure in circulation is a 4.55% character error rate — a vendor-reported number, on a metric used mainly for Chinese, with no independent English word-error-rate benchmark we could find. The model covers 30 languages and 16 Chinese dialects, per the same summary. Treat its accuracy on your accents, your line noise and your jargon as unknown until you measure it.


Where Does Your Audio Actually Go?

If you are outside mainland China, Singapore is the region you would use, and Singapore means Alibaba's "International" inference scope, not a single-country one. Alibaba's deployment-mode documentation says static data stays in the region you choose, while inference runs inside a "service deployment scope"; Singapore and Beijing each have a single fixed scope — International and Mainland China respectively. Regions such as Frankfurt and Virginia let you pick a geographically bounded scope, but the 3.1 ASR model page lists pricing for Beijing and Singapore only.

For a US or EU contact center with customer PII, payment details read aloud, or health information in its recordings, that is the decision that comes before any pilot — and for many regulated buyers it will end the conversation regardless of price. For call-recording archives with a low sensitivity rating, or non-regulated internal meeting audio, it may not.

On training, Alibaba's privacy notice says it "will never use your data for model training." That is a policy statement, not a DPA clause; get it into the contract. The ASR model is hosted-only, with no private deployment option, so self-hosting is not a way around the residency question.

We laid out the residency tests that actually matter in our provider-by-provider residency guide. Apply the same questions here.

What This Does to the Rent-vs-Self-Host Math

At two cents an hour, self-hosting speech recognition stops being a cost decision and becomes a residency and control decision. Our Deepgram vs AssemblyAI vs Whisper analysis put the self-host crossover at volumes where GPU utilisation beats a per-minute API bill. A hosted API at $0.019 an hour pushes that crossover far out — GPU hours, an on-call rotation and model upgrades rarely beat it on cost alone.

The steel-man for staying put is strong. Deepgram and AssemblyAI ship streaming, diarization, redaction and in-region deployments that Qwen's base ASR model does not match today, and your integrations, QA scorecards and redaction pipeline are built on their output formats. Switching costs are real, and a vendor that has served your contact center for three years has earned something.

But the price anchor has moved, and your incumbent knows it. The same happened in text models: once a credible low-price option existed, list prices stopped being the reference point for negotiations, as we argued in our comparison of per-token pricing. You do not need to migrate to Qwen to use this. You need a per-hour number you computed yourself.

On the TTS side, the roughly 70% cut matters for anyone paying ElevenLabs-level rates for IVR prompts or voice agents, but the newest TTS-Next model is priced for Beijing only with a limit of 3 requests per second. For most Western buyers it is a benchmark, not an option.

What to Do Before Your Speech-Vendor Renewal

This Week:

  1. Pull last quarter's transcription invoice and compute your actual cost per audio hour, including diarization, redaction and streaming add-ons. That is the number every comparison starts from.
  2. Classify your audio by sensitivity: which recordings contain PII, payment details or health information, and which could legally be processed in Singapore's International scope. Get your privacy lead to sign the split.

This Month:

  1. Run 50 real calls — your worst line quality, your heaviest accents, your product names — through Qwen-Audio 3.1 ASR in Singapore and through your incumbent. Score word error rate on a hand-corrected sample. Record audio_tokens and seconds from each response to confirm the 25-tokens-per-second rate.
  2. Test chunking on your longest calls: split at silence, stitch the transcripts, and count split words at boundaries. If the stitch is ugly, request filetrans pricing.

Before Renewal:

  1. Bring the per-hour table — theirs, AssemblyAI's, Deepgram's, and Qwen's with the working shown — to the negotiation. Ask for a rate reset or a price-protection clause tied to published list prices.
  2. If you run dual-channel audio, make sure every quote is per channel-hour, not per call-hour. Otherwise you are comparing a doubled number to a single one.

The Bottom Line

Speech recognition is going the way text tokens went in 2024: the capability did not change overnight, the reference price did. Alibaba's cut is real, but it arrives in a unit your procurement team does not use, in a region your privacy team may not approve, with a request cap your architects will have to engineer around.

None of that makes it irrelevant. It makes it a number you should calculate before your vendor does it for you.

Two cents an hour is not a migration plan. It is a negotiating position.

Continue Reading

Share:

Frequently Asked Questions

How much does Qwen-Audio 3.1 speech recognition cost per hour?

About $0.019 per audio hour in Alibaba's Singapore region and about $0.015 in Beijing. Audio is billed at 25 input tokens per second (90,000 per hour) at $0.15 per million in Singapore, plus the transcript as output tokens at $0.47 per million.

How does Qwen-Audio 3.1 ASR pricing compare with Deepgram and AssemblyAI?

As listed on September 27, 2026, AssemblyAI Universal-2 async costs $0.15 per hour and Deepgram Nova-3 pre-recorded costs $0.0043 per minute ($0.258 per hour) pay-as-you-go. Qwen-Audio 3.1 ASR in Singapore works out to roughly one-eighth of AssemblyAI and one-thirteenth of Deepgram, before add-ons.

Is the Qwen-Audio 3.1 price cut really 95%?

Only in the best case. Against the previous Qwen3-ASR-Flash price of $0.000032 per second (about $0.115 per hour, as listed on September 27, 2026), the Singapore rate of about $0.019 per hour is roughly an 84% reduction.

What are the limits of Qwen-Audio 3.1 ASR Flash?

Each request accepts at most 7,168 input tokens, just under five minutes of audio, and returns at most 1,024 output tokens. The model does not support batch inference, fine-tuning or context caching, and is rate-limited to 600 requests per minute. Longer recordings need chunking or the separate filetrans variant.

Where is Qwen-Audio 3.1 audio processed for non-Chinese customers?

The model page lists pricing for Beijing and Singapore. Singapore stores data in-region but runs inference in Alibaba's International deployment scope, which is not a single-country boundary, so regulated buyers should check it against their residency requirements.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →