Smallest.ai
by Smallest.ai
Sub-second voice AI — small, specialised speech models for real-time enterprise voice agents
Smallest.ai builds small, specialised speech models — text-to-speech, speech-to-text, speech-to-speech and a compact language model — plus a platform for deploying real-time voice agents. It targets enterprises in regulated sectors such as financial services, healthcare and contact centres that need interruptible, low-latency phone conversations rather than narration or dubbing, and it competes on speed and cost per minute rather than expressive range.
Smallest.ai is a voice AI company built on the thesis that real-time conversation needs many small specialised models rather than one large general-purpose one. Its stack is four models plus a deployment platform: Lightning, a text-to-speech model quoting sub-100ms time-to-first-audio across 15 or more languages at 44.1kHz; Pulse, speech-to-text covering 38 languages with speaker and emotion detection; Electron, a roughly four-billion-parameter small language model that handles the real-time conversational turn and hands off to a larger offline LLM only when a query exceeds its knowledge; and Hydra, a native speech-to-speech model whose asynchronous architecture runs listening, reasoning, acting and responding in parallel to cut the dead air that makes an interruptible phone call feel wrong. These sit behind Atoms, a dashboard, playground and API for configuring voices, languages and agent behaviour and deploying to rented phone numbers or a web widget. Founded in 2023 by Sudarshan Kamath and Akshat Mandloi, the company employs roughly 60 people and closed a $13 million Series A on 30 July 2026 led by Seligman Ventures with Sierra Ventures and 3one4 Capital participating, taking total funding past $21 million within about nine months of an $8 million seed. It launched Voice 4.0 alongside the round and reports a 76% naturalness win rate against OpenAI's gpt-4o-mini-tts, a 5.38% word error rate and a 3.89 mean opinion score. RingCentral and Truecaller are named customers. The company sells into financial services, healthcare and contact centres with SOC 2 Type 2, ISO 27001, GDPR, HIPAA and PCI compliance plus on-premise and private-cloud deployment, and frames its edge as hardware-level: roughly 550 simultaneous voice calls on about $27,000 of P100 accelerators against around $100,000 of NVIDIA L40S GPUs.
A contact-centre or CX engineering lead in a regulated industry who needs voice agents that can be interrupted mid-sentence and still cost cents per minute at production volume.
Sub-second, interruptible voice conversations at roughly $0.09 to $0.21 per minute, deployable on-premise where the audio is not allowed to leave the network.
At a Glance
- Category
- Audio & Voice
- Pricing
- Usage-based, Freemium, Contact for pricing
- Target Market
- CTOs, VP Customer Experience, Contact Centre Leaders, Enterprise Developers, Product Engineers
- Deployment
- Cloud-first, API-based, Self-hosted, Hybrid
- Founded
- 2023
- Team Size
- 51-200
Key Features
- ✓Lightning text-to-speech
Sub-100ms time-to-first-audio at 44.1kHz across 15 or more languages, with prosody tuned for conversational turn-taking rather than narration.
- ✓Hydra speech-to-speech
Asynchronous architecture that listens, reasons, acts and responds in parallel, cutting the lag that makes interruptions feel unnatural on a call.
- ✓Pulse speech-to-text
Transcription across 38 languages with speaker and emotion detection, so an agent can react to tone as well as to words.
- ✓Electron small language model
A roughly four-billion-parameter model handling the real-time turn, escalating to a larger offline LLM only for genuinely complex queries.
- ✓Atoms agent platform
Dashboard, playground and API for configuring voices, languages and behaviour, then deploying to phone numbers or a web widget.
- ✓Fast voice cloning
Production voice cloning from as little as five seconds of reference audio, avoiding the studio recording sessions rivals require.
- ✓On-premise and private-cloud deployment
The full model stack runs inside the customer's own environment, which is what makes regulated healthcare and banking deployments possible at all.
Capabilities
Use Cases
- •Contact-centre call deflection
A bank routes routine balance and payment queries to a voice agent that callers can interrupt without the conversation falling apart.
- •Outbound appointment and collections calling
Healthcare or lending teams run high-volume outbound campaigns where per-minute model cost decides whether the unit economics work.
- •Multilingual customer support
A support line handles callers who code-switch mid-sentence, with the agent following the language change rather than losing the thread.
- •Embedding speech into an existing product
A software vendor adds voice input and output to its own application through the API without building a speech stack in-house.
- •On-premise voice for regulated data
A hospital deploys the models inside its own infrastructure so patient audio never crosses the network boundary to a third party.
Ideal For
Best For
- ✓Real-time inbound and outbound phone agents in contact centres where response lag breaks the conversation
- ✓Regulated deployments in healthcare and financial services needing HIPAA and PCI compliance with on-premise or private-cloud hosting
- ✓High-volume voice workloads where per-minute cost rather than expressive range is the binding constraint
- ✓Multilingual voice agents needing mid-sentence language switching across 15 languages
- ✓Product teams embedding text-to-speech or speech-to-text into an existing application via API rather than buying a whole agent platform
Not Ideal For
- ✗Audiobook, narration, dubbing and podcast production — a competitor's head-to-head comparison describes the voices as lacking depth and emotional range, and professional voice cloning is not supported
- ✗Buyers who want the largest voice library and deepest customisation, since the same comparison rates the feature set as more limited and pronunciation context weaker than ElevenLabs
- ✗Teams optimising for the lowest published latency alone, where Cartesia quotes roughly 40 to 90ms against Smallest's sub-100ms plus network time
Integrations
Deployment
Market Analysis
Pros
- ✓Genuinely fast, with sub-100ms time-to-first-audio and an architecture built around interruption, which is the failure mode that makes most voice agents unusable on a phone call
- ✓Per-minute pricing of $0.09 to $0.21 is low enough that high-volume calling economics work, and the component breakdown is published rather than buried
- ✓Unusually complete compliance for a company founded in 2023 — SOC 2 Type 2, ISO 27001, HIPAA, PCI and GDPR — with on-premise deployment available
- ✓Named production customers in RingCentral and Truecaller, plus published benchmark figures (76% naturalness win rate over gpt-4o-mini-tts, 5.38% word error rate, 3.89 mean opinion score) rather than vague marketing claims
- ✓Deliberately narrow scope: it does one thing, real-time conversational voice, instead of spreading across dubbing and content generation
Cons
- ✗A competitor's head-to-head comparison rates the voices as lacking depth and emotional range, with less contextual awareness in pronunciation than ElevenLabs
- ✗Professional voice cloning is not supported and customisation options are characterised as basic relative to the market leader
- ✗The latency claims are model-only figures measured under favourable conditions and exclude network time, and Cartesia quotes lower numbers for its own model
- ✗Independent practitioner coverage is thin — no G2, Capterra or substantive Hacker News discussion surfaced, so there is very little unfiltered production experience to read before buying
- ✗Self-serve concurrency is capped at 20 simultaneous calls, so any real contact-centre deployment is an Enterprise negotiation with unpublished pricing
Pricing
Pay As You Go
From $0.09/minute
- ✓$10 in free credits to start
- ✓Unlimited agents
- ✓Full API suite access
- ✓20 concurrent calls
- ✓Web widget
- ✓Basic testing suite
- ✓Post-call analytics (10 per agent)
- ✓Email and community support
Enterprise
Contact for pricing
- ✓Dedicated infrastructure with 99.99% uptime SLA
- ✓On-premise deployment
- ✓SSO and advanced compliance (SOC 2, HIPAA)
- ✓Dedicated forward-deployed engineers
- ✓Priority and prompt-engineering support
- ✓Custom integrations and branding
Metered per minute of conversation rather than per seat: $0.09 to $0.21 all-in depending on model choice, and unusually the vendor publishes the component breakdown — roughly $0.009 per minute for speech-to-text, $0.09 for text-to-speech, $0.01 flat for hosting, plus whichever LLM you route to at $0.045 to $0.10 per minute. Phone numbers rent at $10 per month each and text messages cost $0.005. New accounts get $10 in free credits. Everything that makes it deployable inside a bank or hospital — on-premise hosting, SSO, the 99.99% SLA, dedicated infrastructure — is Enterprise-only and unpriced, and the self-serve tier caps concurrency at 20 calls, so real contact-centre volume forces a sales conversation.
Security & Compliance
Connect
Sources
This page was written from 5 sources, 3 on domains other than smallest.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Ringg AI
Multilingual voice, chat and WhatsApp agents for high-volume enterprise conversations
Rime
Text-to-speech trained on studio-recorded conversation, built for regulated call volume
Vapi
Voice AI infrastructure for developers who want control, not a packaged agent
ElevenLabs
Production-grade voice AI: speech, cloning, dubbing and deployable voice agents in 70+ languages.