AssemblyAI
by AssemblyAI
Speech-to-text, speech understanding and voice agent APIs priced per hour of audio
AssemblyAI is a speech AI API company that sells pre-recorded and real-time transcription, speech understanding add-ons and an all-in-one Voice Agent API. It is built for product and engineering teams adding transcription, call analytics or voice agents to their software without training their own speech models.
AssemblyAI is a San Francisco company founded in 2017 by Dylan Fox that builds its own speech recognition models and sells them through developer APIs. Its flagship models, Universal-Async 3.5 and Universal-Realtime 3.5, shipped on July 7, 2026, expanding coverage from 6 to 18 languages with mid-sentence code-switching; the company reports word error rate falling from 9.07% to 7.69% on its benchmark set. July 2026 also brought a Sync API that returns finished transcripts for clips up to two minutes in a single HTTP call at roughly 134 ms median latency, plus in-house text-to-speech voices inside the Voice Agent API. In August it added entity-aware endpointing for voice agents and a 1.0 Python SDK covering async, streaming and sync transcription. The product line now spans pre-recorded and streaming speech-to-text, Speech Understanding add-ons (speaker identification, entity detection, PII redaction, translation, summarisation), Guardrails, an LLM Gateway that routes to Claude, GPT and Gemini models, and the Voice Agent API at $4.50 per hour bundling STT, LLM, TTS and orchestration. Pricing is public and billed per second. Compliance covers SOC 2 Type 2, ISO 27001, GDPR, PCI-DSS for the Voice Agent API and a HIPAA BAA signable without a sales call, with an EU region at US prices. The company raised a $50 million Series C led by Accel in December 2023 and lists Insight Partners and Y Combinator among its backers. It competes with Deepgram, ElevenLabs, OpenAI's speech models and the speech services of AWS, Google and Azure.
An engineering lead building call analytics, meeting transcription or voice agents who wants a published per-hour rate and a HIPAA BAA without a sales cycle.
Accurate multilingual transcription and speech analytics through one API, with self-serve pricing from $0.15 per audio hour.
At a Glance
- Category
- Audio & Voice
- Pricing
- Usage-based, Freemium
- Target Market
- CTOs, Enterprise Developers, Product Managers, Contact Center Leaders
- Deployment
- API-based, Cloud-only
- Founded
- 2017
- Headquarters
- San Francisco, United States
- Team Size
- 51-200
- Customers
- Hundreds of enterprise customers, including dozens of Fortune 500 companies (vendor figure)
Key Features
- ✓Universal-3.5 speech models
Async and real-time models covering 18 languages with mid-sentence code-switching and a reported 7.69% word error rate.
- ✓Sync STT API
Returns a finished transcript for clips up to two minutes in one HTTP call at about 134 ms median latency, which suits dictation and short commands.
- ✓Voice Agent API
Bundles speech-to-text, an LLM, text-to-speech and turn-taking for $4.50 per hour, so a team can ship a voice agent without stitching vendors together.
- ✓Speech Understanding add-ons
Speaker identification, entity detection, PII redaction, translation and summarisation billed per hour and stacked on any transcript.
- ✓Entity-aware endpointing
Stops voice agents cutting a caller off mid phone number or email address; the vendor reports a 40.7% gain on phone-number endpointing.
- ✓LLM Gateway
Routes transcripts to Claude, GPT and Gemini models at per-token rates with prompt caching, keeping the speech and LLM bill on one account.
- ✓EU data residency and HIPAA BAA
An EU endpoint at US prices and a BAA signable in minutes let regulated teams start without a procurement negotiation.
Capabilities
Use Cases
- •Call-centre quality analytics
Transcribe every support call, redact card numbers and names, and extract entities so QA teams can review far more calls than manual sampling allows.
- •Clinical documentation
Use Medical Mode and a signed HIPAA BAA to transcribe patient conversations for downstream note generation in healthcare software.
- •Phone-based voice agents
Run an inbound phone agent on the Voice Agent API with entity-aware endpointing so callers can read out phone numbers without being interrupted.
- •Meeting intelligence products
Transcribe recorded meetings with speaker identification and summarisation to power searchable notes and action items inside a SaaS product.
Ideal For
Best For
- ✓Contact-centre call analytics with PII redaction and entity detection
- ✓Meeting and interview transcription with speaker identification
- ✓Building real-time voice agents on a single bundled API
- ✓Healthcare transcription that needs a HIPAA BAA and Medical Mode
- ✓EU workloads that must stay in-region for GDPR
Not Ideal For
- ✗Buyers who need hard spending caps on a self-serve account; there are balance alerts and auto-recharge, but nothing stops a runaway batch job mid-run
- ✗Teams that must run speech models fully on-device or inside an air-gapped network
- ✗Non-English-first products that need published per-language accuracy benchmarks, since AssemblyAI's benchmarks remain English-focused
Integrations
Deployment
Market & Ratings
Hundreds of enterprise customers, including dozens of Fortune 500 companies (vendor figure)
Market Analysis
Pros
- ✓Transparent per-second pricing with $50 of free credits and no card required
- ✓Broad compliance: SOC 2 Type 2, ISO 27001, GDPR, PCI-DSS and a self-serve HIPAA BAA
- ✓Fast release cadence in 2026: Universal-3.5, Sync API, in-house TTS and entity-aware endpointing within two months
- ✓EU data residency at the same price as the US region
Cons
- ✗Stacked add-ons and the September 2026 effort-based rates make the real bill hard to estimate without a calculator
- ✗No hard spend caps on self-serve accounts, so a runaway job can overspend
- ✗Hacker News users reported speaker diarization still needing manual confirmation of who is who
- ✗Voice Agent API is a bundle with no published a-la-carte price for its LLM and TTS parts
Pricing
Free
$0
- ✓$50 in free credits on signup
- ✓No credit card required
- ✓Streaming limited to 5 new streams per minute
Pay-as-you-go
From $0.15/hr of audio
- ✓Universal-2 async $0.15/hr
- ✓Universal-3.5 Pro async $0.21/hr
- ✓Streaming $0.15-$0.45/hr
- ✓Voice Agent API $4.50/hr
- ✓100 new streams per minute
Enterprise
Contact for pricing
- ✓Volume discounts
- ✓Dedicated support
- ✓Custom contract pricing
Billing is per second of audio with no rounding, from $0.15/hr (Universal-2) and $0.21/hr (Universal-3.5 Pro). Add-ons stack: a typical call-analytics pipeline with PII redaction, entity detection and translation runs about $0.43/hr effective by one independent estimate. Since September 20, 2026 some add-ons bill at two rates depending on a request-time effort setting. Enterprise volume discounts are negotiated and unpublished. Prices checked on assemblyai.com/pricing on October 2, 2026.
Security & Compliance
Sources
This page was written from 6 sources, 5 on domains other than assemblyai.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Cresta
Contact center AI platform for autonomous AI agents, real-time agent assist and conversation intelligence
Deepgram
Voice AI API platform for real-time speech-to-text, text-to-speech and voice agents
PolyAI
Enterprise voice AI agents for customer service, built on a proprietary dialog model
Cartesia
Low-latency voice AI on state space models: Sonic text-to-speech, Ink speech-to-text and managed voice agents