Cartesia
by Cartesia
Low-latency voice AI on state space models: Sonic text-to-speech, Ink speech-to-text and managed voice agents
Cartesia is a real-time voice AI platform offering Sonic text-to-speech, Ink streaming speech-to-text and managed voice agents, built on state space models rather than transformers. It is for contact-center, voice-agent and embedded-device teams where time-to-first-audio is user-facing and deployment must span cloud, VPC or on-device.
Cartesia is a voice AI company founded in 2023 by Stanford AI Lab researchers Karan Goel (CEO), Albert Gu, Arjun Desai and Brandon Yang, with Stanford professor Chris Ré on the founding team; Gu co-created the S4 and Mamba state space model (SSM) architectures, which Cartesia uses instead of transformers to get low latency and efficient long-context streaming. Its product line has three parts: the Sonic family of text-to-speech models, the Ink streaming speech-to-text models, and a managed voice-agent platform with telephony billed per minute. Sonic-3, launched alongside a $100M round from Kleiner Perkins, Index Ventures, Lightspeed and NVIDIA announced in late 2025, added laughter, emotion and tonal control with a reported 90 ms model latency and 190 ms end-to-end, across 42 languages. The cadence has been fast: Sonic-3.5 took first place on Artificial Analysis's speech leaderboard in May 2026, and Sonic-3.6 entered beta on 17 August 2026 with 44 languages, better Hinglish and natural disfluencies without SSML, again ranking first on that leaderboard as of 19 August 2026. Cartesia has raised roughly $191M in total ($27M seed December 2024, $64M Series A March 2025, $100M late 2025). Named customers include ServiceNow, Decagon, Cresta, Quora, Retell AI and Zomato, and the company reported more than 10,000 customers at its Series A. It states SOC 2 Type II, HIPAA, PCI-DSS and GDPR compliance with optional zero data retention, and deploys via cloud, VPC/on-premises or on-device. It competes with ElevenLabs, Deepgram, OpenAI's Realtime API and voice-agent platforms such as Vapi and Retell.
Heads of contact-center AI and voice-agent platform engineering who need the lowest time-to-first-audio and compliant (HIPAA, PCI) deployment options.
Leaderboard-ranked, natural-sounding speech at roughly 90 ms model latency, so phone agents stop sounding slow and robotic.
At a Glance
- Category
- Audio & Voice
- Pricing
- Freemium, Subscription, Usage-based
- Target Market
- CTOs, Heads of Customer Experience, Enterprise Developers, Voice AI Product Teams
- Deployment
- API-based, Cloud-first, Self-hosted, Edge-first
- Founded
- 2023
- Customers
- 10,000+ (reported at Series A, March 2025)
Key Features
- ✓Sonic text-to-speech
SSM-based TTS with emotion, laughter and natural disfluencies across 44 languages, ranked first on Artificial Analysis in August 2026.
- ✓Ink speech-to-text
Streaming transcription designed for conversational turn-taking, so a single vendor covers both ends of a voice pipeline.
- ✓Managed voice agents
Hosted agent platform with provisioned phone numbers, billed at $0.06 per call minute plus $0.014 per minute telephony.
- ✓Voice cloning and localization
Instant and professional voice cloning plus accent localization, letting brands keep one consistent voice across markets.
- ✓Flexible deployment
Cloud with regional endpoints, VPC/on-premises and on-device options for latency, residency and embedded or robotics use.
- ✓Compliance posture
Vendor states SOC 2 Type II, HIPAA, PCI-DSS service provider and GDPR, with optional zero data retention for sensitive calls.
Capabilities
Use Cases
- •Contact-center automation
Customer-service agents answer calls with sub-200 ms responses, reducing awkward silences that cause callers to talk over the bot.
- •Voice-agent platform backend
Agent vendors plug Sonic and Ink into their orchestration layer to power millions of monthly conversations for their own customers.
- •Healthcare patient outreach
Providers run scheduling and reminder calls under a BAA with zero data retention to keep protected health information out of logs.
- •Multilingual customer support
Brands serve callers in 44 languages, including Hinglish, using a cloned brand voice localized to regional accents.
Ideal For
Best For
- ✓Real-time phone and contact-center voice agents where response latency is audible to callers
- ✓Voice-agent platforms (the Retell/Decagon pattern) needing a fast TTS and STT pair from one vendor
- ✓Regulated voice workflows in healthcare and payments that require HIPAA BAAs or PCI-DSS
- ✓On-device or VPC voice for robotics, embedded products or data-residency constraints
- ✓Multilingual deployments across 40+ languages, including Hinglish for Indian markets
Not Ideal For
- ✗Audiobook, character-voice and long-form narration production — third-party testing finds ElevenLabs stronger there, with a far larger voice library
- ✗Teams that require open model weights to self-modify or fine-tune
- ✗Builders committed to end-to-end speech-to-speech models rather than a modular STT-LLM-TTS pipeline
- ✗Production teams unable to absorb frequent model upgrades — three major Sonic versions shipped within about a year, adding regression-testing burden
Integrations
Deployment
Market & Ratings
10,000+ (reported at Series A, March 2025)
Market Analysis
Pros
- ✓Top-ranked TTS quality on Artificial Analysis while keeping latency around 90 ms model / 190 ms end-to-end
- ✓Cheap self-serve entry with a free tier and a $5/month commercial plan
- ✓Strong compliance story for voice (SOC 2 Type II, HIPAA, PCI-DSS, GDPR, zero data retention)
- ✓Well funded (~$191M) with enterprise customers such as ServiceNow and Decagon
Cons
- ✗Trails ElevenLabs on long-form narration and character voices, with ~100 curated voices vs 1,000+
- ✗Credit-based pricing obscures true unit cost and concurrency caps force tier upgrades
- ✗Managed agent platform is young and less mature on telephony than Vapi, Retell or LiveKit
- ✗Rapid model churn (Sonic 2, 3, 3.5, 3.6 in about 18 months) adds regression-testing work for production teams
- ✗Thin independent community signal — HN launch posts drew few comments, so evaluation leans on benchmark sites
Pricing
Free
$0
- ✓20K credits/month
- ✓TTS and STT
- ✓2 concurrent TTS / 8 STT
- ✓1 phone number
Pro
From $5/mo
- ✓100K credits (~133 TTS minutes)
- ✓Commercial license
- ✓Instant voice cloning
Startup
From $49/mo
- ✓1.25M credits (~1,667 TTS minutes)
- ✓2 professional voice clones
- ✓Organizations
Scale
From $299/mo
- ✓8M credits (~10,667 TTS minutes)
- ✓Priority support
- ✓15 concurrent TTS / 60 STT
Enterprise
Contact for pricing
- ✓Volume pricing
- ✓Custom concurrency
- ✓DPAs and BAAs
- ✓SSO
- ✓Shared Slack channel
List prices are published but metered in credits (about 750 credits per TTS minute on self-serve tiers), which obscures unit cost; concurrency caps per tier can force upgrades independent of volume. Voice-agent calls add $0.06/min plus $0.014/min telephony. BAAs, SSO and custom concurrency are Enterprise-only.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 5 on domains other than cartesia.ai.
- 1.cartesia.ai — pricingvendor
- 2.cartesia.ai — cartesia.aivendor
- 3.cartesia.ai — gdpr compliancevendor
- 4.fortune.com — exclusive cartesia voice ai startup raises 64 million series
- 5.slator.com — cartesia launches text to speech model
- 6.aivoicenewsletter.com — cartesia s 100m sonic 3 leap
- 7.rywalker.com — cartesia
- 8.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Omilia Cloud Platform
Self-learning agentic voice AI for enterprise contact centers: voice agents, voice biometrics and fraud defense on one platform
Phonely
Voice agents running on Alma, a model trained on ten million real phone calls
Smallest.ai
Sub-second voice AI — small, specialised speech models for real-time enterprise voice agents
Ringg AI
Multilingual voice, chat and WhatsApp agents for high-volume enterprise conversations