Deepgram
by Deepgram
Voice AI API platform for real-time speech-to-text, text-to-speech and voice agents
Deepgram is a voice AI API platform that gives developers and enterprises real-time speech-to-text (Nova-3, Flux), text-to-speech (Aura-2) and a unified Voice Agent API. It is built for teams shipping voice agents, contact-centre products and live captioning who need low latency, usage-based pricing and the option to self-host.
Deepgram is a San Francisco voice AI company, founded in 2015 and led by co-founder and CEO Scott Stephenson, that sells speech models as APIs rather than finished applications. Its catalogue covers speech-to-text with the Nova-3 family (monolingual, multilingual and medical variants), Flux, text-to-speech with Aura-2, and a Voice Agent API that combines speech recognition, LLM orchestration and speech synthesis in one real-time service. Flux, launched in October 2025, is described by Deepgram as its first Conversational Speech Recognition model: it fuses transcription and end-of-turn detection in a single model so a voice agent knows when a caller has finished speaking from context rather than silence, which Deepgram says reduces false interruptions by roughly 30% and cuts agent response latency by 200-600 ms versus pipeline approaches; at launch it was English-only. Deployment options include Deepgram's cloud, self-hosted Docker/Podman or Kubernetes (on the Enterprise plan) and Amazon SageMaker via AWS Marketplace. In January 2026 Deepgram raised a $130M Series C at a $1.3B valuation led by AVP, with strategic participation from Twilio, ServiceNow Ventures, SAP and Citi Ventures, and acquired OfOne, a voice-ordering platform for quick-service restaurants. The company reports 1,300+ organisations and 200,000+ developers on the platform, and Sacra reports it reached $100M ARR in August 2026, with voice-agent companies such as Decagon, Sierra and Vapi among its customers. It competes with AssemblyAI, Speechmatics, Gladia, ElevenLabs and open-weight Whisper deployments, generally positioned on speed and price rather than raw benchmark accuracy.
Engineering leaders building real-time voice agents or contact-centre products, who need low-latency STT/TTS priced per minute with a self-host path for regulated data.
One API vendor for listening, turn-taking and speaking in a voice agent, with published per-minute rates and turn detection built into the recognition model.
At a Glance
- Category
- Audio & Voice
- Pricing
- Usage-based, Freemium
- Target Market
- CTOs, Enterprise Developers, Heads of Customer Experience, AI Engineers
- Deployment
- API-based, Cloud-first, Self-hosted
- Founded
- 2015
- Headquarters
- San Francisco, United States
- Customers
- 1,300+ organisations; 200,000+ developers (Jan 2026)
Key Features
- ✓Nova-3 speech-to-text
Streaming and pre-recorded transcription in monolingual, multilingual and medical variants, with keyterm prompting for domain vocabulary.
- ✓Flux conversational speech recognition
Fuses transcription and context-aware end-of-turn detection in one model, so voice agents respond faster and interrupt callers less often.
- ✓Aura-2 text-to-speech
Low-latency speech synthesis priced per thousand characters, designed to pair with Deepgram STT inside real-time voice agents.
- ✓Voice Agent API
A single speech-to-speech endpoint combining STT, LLM orchestration and TTS, billed per minute, removing the need to stitch three vendors together.
- ✓Self-hosted deployment
Enterprise customers can run Deepgram models on their own GPUs via Docker/Podman or Kubernetes, or through Amazon SageMaker, keeping audio inside their perimeter.
- ✓Compliance controls
SOC 2 Type 1 and 2, GDPR compliance with EU data residency, and HIPAA BAAs make it usable for regulated audio such as patient calls.
Capabilities
Use Cases
- •Customer-service voice agents
Power an AI phone agent with Flux turn detection and Aura-2 voices so it answers promptly without cutting callers off.
- •Contact-centre analytics
Transcribe every inbound and outbound call in real time to feed QA scoring, compliance review and agent coaching dashboards.
- •Clinical and regulated transcription
Use Nova-3 Medical under a HIPAA BAA or self-hosted deployment to transcribe patient conversations without audio leaving approved infrastructure.
- •Restaurant drive-thru ordering
Automate quick-service drive-thru and phone orders using the OfOne-derived Deepgram for Restaurants offering, which reports over 95% containment.
- •Live captions for media and meetings
Stream low-latency captions for broadcasts, webinars and internal meetings using the streaming Nova-3 endpoint.
Ideal For
Best For
- ✓Real-time voice agents that need accurate end-of-turn detection to avoid talking over callers
- ✓Contact-centre and call-analytics products transcribing high volumes of streaming audio
- ✓Live captioning and meeting transcription where latency matters more than batch accuracy
- ✓Regulated teams (healthcare, finance) that need HIPAA BAA, EU data residency or self-hosted models
- ✓Voice ordering and drive-thru automation, via the OfOne-based restaurant offering
Not Ideal For
- ✗Teams needing only occasional one-off transcription — the model choice (Nova-3 vs Flux vs Medical) and API-first setup are more than they need
- ✗Workloads where top batch accuracy and entity capture (names, emails, spelled sequences) matter most — independent comparisons put AssemblyAI ahead on WER for read speech
- ✗Heavily multilingual or code-switching products — Flux launched English-only and Nova-3 code-switching covers a limited set of languages
- ✗Buyers wanting built-in translation or entity detection as STT add-ons, which Deepgram does not offer directly
Integrations
Deployment
Market & Ratings
1,300+ organisations; 200,000+ developers (Jan 2026)
Market Analysis
Pros
- ✓Strong real-time streaming speed; an independent review scored it 9.9/10 on speed
- ✓Flux turn detection addresses the main UX failure of voice agents — interrupting or lagging the caller
- ✓Transparent, low per-minute list pricing with $200 of free credit
- ✓Self-hosted, SageMaker and EU-residency options for regulated buyers
Cons
- ✗Independent and competitor benchmarks put AssemblyAI ahead on accuracy for read speech (reported 5.93% vs 7.9% average WER) and on entity capture
- ✗Users report inconsistencies with accents and exact entities such as email addresses, names and spelled-out sequences
- ✗Hacker News practitioners flag limited real-time diarization for larger groups compared with rivals like Soniox
- ✗Model selection (Nova-3, Multilingual, Medical, Flux) is not interchangeable and requires deliberate testing; Flux launched English-only
Pricing
Pay As You Go
$0 to start ($200 free credit)
- ✓No credit card required
- ✓No minimums or expiration
- ✓Nova-3 streaming from $0.0048/min
- ✓Flux English $0.0065/min
- ✓Aura-2 $0.030 per 1K characters
- ✓Voice Agent API from $0.075/min
Growth
From $4,000/yr (prepaid credits)
- ✓Up to 20% lower per-unit rates
- ✓Annual prepaid credits redeemed against usage
Enterprise
Contact for pricing
- ✓Custom volume pricing
- ✓Self-hosted deployment
- ✓HIPAA BAA
- ✓EU data residency
Deepgram publishes list prices: speech-to-text is metered per audio minute (Nova-3 streaming $0.0048/min, Flux $0.0065/min on pay-as-you-go), text-to-speech per 1,000 characters, and the Voice Agent API per minute ($0.075 Standard, $0.163 Advanced). New accounts get $200 credit. Growth prepay (min $4K/yr) saves up to 20%; self-hosting requires the Enterprise plan. Competitor reviews note add-on features are billed separately, which complicates cost estimates.
Security & Compliance
Connect
Sources
This page was written from 9 sources, 6 on domains other than deepgram.com.
- 1.deepgram.com — press release deepgram raises series cvendor
- 2.deepgram.com — pricingvendor
- 3.deepgram.com — fluxvendor
- 4.developers.deepgram.com — self hosted introduction
- 5.builtinsf.com — deepgram raises 130m series c 1b valuation 20260115
- 6.sacra.com — deepgram
- 7.diyai.io — deepgram ai review
- 8.gladia.io — assemblyai vs deepgram
- 9.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
PolyAI
Enterprise voice AI agents for customer service, built on a proprietary dialog model
Cartesia
Low-latency voice AI on state space models: Sonic text-to-speech, Ink speech-to-text and managed voice agents
Omilia Cloud Platform
Self-learning agentic voice AI for enterprise contact centers: voice agents, voice biometrics and fraud defense on one platform
Phonely
Voice agents running on Alma, a model trained on ten million real phone calls