Lowest-latency voice model on the market — 90ms from text to audio. Powers real-time voice agents and call-center bots that don't feel robotic.
Audio & Voice
Cartesia Sonic is a text-to-speech model built for one job: pushing the first byte of audio out the door before your caller notices the model is thinking. The current model, Sonic 3.5, streams first-token audio in roughly 90 milliseconds. That is not a marketing number. It is the reason Cartesia exists as a company distinct from ElevenLabs, PlayHT, and OpenAI's Realtime API — every other realistic-voice vendor spent 2023 to 2025 optimizing for naturalness first and latency second, and Cartesia inverted the priority.
The honest pitch: if you are shipping a voice agent that talks to a human on a phone line, in a browser, or through a headset, Sonic is likely the fastest realistic option you can hit from an API today. Time-to-first-audio is the number that decides whether your bot sounds like a bot. Ninety milliseconds is under the threshold where humans register a "gap" in conversational rhythm. ElevenLabs Flash v2.5 quotes about 75 milliseconds for the same benchmark, so Cartesia no longer owns the crown outright — but Sonic still edges Flash on naturalness in third-party arenas (it currently ranks #1 on Artificial Analysis Speech Arena) while Flash trades realism for speed. That is the specific pocket Cartesia occupies.
Where it sits on the map: ElevenLabs is the incumbent for realism and voice library depth. OpenAI's Realtime API (gpt-realtime) is the easiest path if you are already deep in OpenAI's stack and want speech-to-speech in one call, but its audio-output pricing works out expensive at scale and the voice roster is narrow. Deepgram Aura-2 undercuts everyone on raw per-character cost and pairs naturally with Deepgram STT for full-stack agent builds. PlayHT sits adjacent to Cartesia on the "voice agent" positioning but its Play 3.0 Mini has not caught Sonic on latency. Cartesia's wedge is the intersection of sub-100ms latency, commercial-grade realism, and enterprise deployment options — on-premise, VPC, and in-region inference. If any of those three matter, it belongs on your shortlist. If none do, cheaper or more established options will serve you fine.
90ms time-to-first-audio. Sonic 3.5 streams the first byte of audio in approximately 90 milliseconds after receiving text. This is the number that matters for voice agents — total generation speed is largely irrelevant once streaming starts, because the user hears audio while the rest of the sentence is still being synthesized. The 90ms figure comes from Cartesia's own benchmarks, which is standard for the category, but it has been reproduced by third parties enough times that it is a fair working number.
Instant voice cloning from 10 seconds of audio. Upload a 10-second clean sample and Sonic produces a usable clone. Not studio-quality — for that you want the professional cloning path, which is a longer intake process gated to the Startup tier and up. The 10-second instant clone is the workflow you actually use in production for prototypes, one-off characters, or user-generated voices. Speaker similarity is high enough that it survives a demo. It is not high enough that a family member would fail to notice.
Emotional calibration from transcript context. Sonic reads the text you send it and infers whether the delivery should be excited, sober, apologetic, or matter-of-fact. You do not have to add SSML emotion tags. This is the feature ElevenLabs eleven_v3 also pushes hard, and it is the correct evaluation axis for post-2025 TTS. Sonic is not the top of that leaderboard for long-form emotive narration — v3 is better at three-paragraph audiobook passages — but for single-sentence agent turns Sonic is competitive and faster to synthesize.
Non-verbal expressions inline. You can insert laughter, sighs, breaths, and other paralinguistic marks directly in the transcript and Sonic renders them. This is table stakes in 2026 but the implementation is cleaner than most, and the marks do not spike latency the way SSML-heavy inputs sometimes do on competing systems.
42-language multilingual coverage. Sonic 3.5 supports around 40 languages natively, and voice clones can be localized into 42. That is short of ElevenLabs' eleven_v3 (70+) and Scribe (90+), so if your product has heavy coverage requirements across Southeast Asian or African languages, Cartesia will have gaps. For the major commercial languages — English, Spanish, Mandarin, Hindi, Arabic, Portuguese, French, German, Japanese — coverage is solid.
Voice Agent runtime, not just TTS. Cartesia sells the underlying Sonic TTS but also a full Voice Agent product that bundles STT, turn-taking, and telephony. Voice Agent calls are billed at $0.06 per minute on every tier, with an additional $0.014 per minute if you use a Cartesia-provided phone number. This matters because Sonic is often bought alongside LiveKit, Vapi, or Retell — you can BYO orchestration, or you can buy the whole stack from Cartesia. Pick your integration boundary deliberately.
Enterprise deployment: on-prem, VPC, in-region. Cartesia will run Sonic in your cloud, on your metal, or in a specific region for compliance. HIPAA, SOC 2 Type 2, GDPR, and PCI covered on the enterprise tier. This is the feature that gets Cartesia into healthcare, finance, and government pilots where ElevenLabs' cloud-only footprint is a non-starter. It is also gated behind an enterprise contract, so it is invisible in the self-serve tiers.
Custom pronunciation dictionaries. You can define how the model pronounces proper nouns, product names, and domain jargon. This is unglamorous but decisive for call-center bots — no one wants "ServiceNow" pronounced as two random words in front of a customer.
Cartesia sells credits, not characters. That is the first thing to know, because it makes cross-vendor math harder than it should be. The tiers, as of August 2026:
| Tier | Price/mo | Credits/mo | Voice Cloning | Commercial Use |
|---|---|---|---|---|
| Free | $0 | 20,000 | None | Not allowed |
| Pro | $5 | 100,000 | Instant | Yes |
| Startup | $49 | 1,250,000 | Instant + Professional | Yes |
| Scale | $299 | 8,000,000 | Instant + Professional | Yes |
| Enterprise | Custom | Custom | All + custom | Yes |
Voice Agent add-on is $0.06 per minute on every tier, plus $0.014 per minute if you rent a Cartesia phone number. Prepaid Agent Minutes are bundled at each tier ($1 free, $5 Pro, $49 Startup, $299 Scale).
The Free tier explicitly forbids commercial use, which is worth naming out loud — you cannot ship a product on it. If you are evaluating for anything past personal projects, budget for the $5 Pro tier at minimum. Professional voice cloning (the studio-grade intake) is gated to the $49 Startup tier and up; the $5 Pro tier only unlocks the 10-second instant clone.
Versus alternatives on raw dollars: ElevenLabs Starter is $5, Creator $22, Pro $99, Scale $299, Business $990. Deepgram Aura-2 is $0.030 per 1,000 characters pay-as-you-go, or $0.027 on Growth — genuinely the cheapest per-character option in the category. OpenAI's gpt-realtime costs $32 per million input audio tokens and $64 per million output, which works out roughly comparable to Cartesia and ElevenLabs on a per-minute basis but bakes in the LLM cost too. PlayHT Creator sits around $39/mo and Unlimited around $99/mo [verify pricing]. On the low end Cartesia and ElevenLabs are effectively tied; on the middle tiers Cartesia is meaningfully cheaper than ElevenLabs; at enterprise the negotiation matters more than the sticker.
Pros
Cons
ElevenLabs — Free tier at $0, Starter $5, Creator $22, Pro $99, Scale $299, Business $990. Pick ElevenLabs if you need the deepest voice library on the market, 70+ language coverage via eleven_v3, or the best-in-class long-form emotive narration for audiobooks and video. Pick Cartesia if you need sub-100ms latency for real-time agents — Flash v2.5 gets close (~75ms) but sacrifices realism to do it, and eleven_v3 is not built for that latency envelope.
OpenAI Realtime API (gpt-realtime) — $32 per million audio input tokens, $64 per million output. Pick OpenAI if you want a single API call for speech-to-speech (STT + LLM + TTS in one round trip) and are already committed to their stack. Pick Cartesia if you want to control your own LLM choice, need voice cloning, or need enterprise deployment options OpenAI does not sell.
Deepgram Aura-2 — $0.030 per 1,000 characters pay-as-you-go, $0.027 on Growth. Pick Deepgram if raw per-character cost dominates your P&L and you already use Deepgram STT — the full agent stack from one vendor is a real advantage. Pick Cartesia if realism and voice cloning matter, because Aura-2 is competent but not at the top of the naturalness leaderboard.
PlayHT — Roughly $39/mo Creator and $99/mo Unlimited [verify pricing]. Pick PlayHT if you want a large stock voice library with commercial licensing baked in and have latency headroom. Pick Cartesia if you need the fastest available time-to-first-audio without giving up realism.
Buy Cartesia Sonic if you are shipping a voice agent — call center, in-app assistant, phone bot, real-time avatar — where sub-100ms time-to-first-audio is the difference between "sounds human" and "sounds like a bot." The $5 Pro tier is the correct entry point for any commercial project because Free forbids commercial use. Move to Startup ($49) the moment you need professional voice cloning or hit 100K credits. Scale ($299) is for teams with real production traffic. Skip Cartesia if your workload is long-form narration (ElevenLabs v3 is better), if you need 70+ languages (ElevenLabs Scribe or v3), if raw per-character cost is the only thing that matters (Deepgram Aura-2), or if you want a single-vendor speech-to-speech stack with your LLM baked in (OpenAI Realtime). Pick Cartesia when the milliseconds decide the product; skip when they don't.
Q: Is Cartesia Sonic free? A: Cartesia offers a Free tier at $0/month with 20,000 credits, but the Free tier explicitly forbids commercial use — you cannot ship a product on it. For any commercial project, the entry tier is Pro at $5/month with 100,000 credits and instant voice cloning.
Q: How does Cartesia Sonic compare to ElevenLabs Flash v2.5 on latency? A: Sonic 3.5 streams first-byte audio in about 90ms; ElevenLabs Flash v2.5 quotes about 75ms. Flash edges Sonic on the pure latency number, but Sonic sits ahead on naturalness — currently #1 on Artificial Analysis Speech Arena — so the trade is speed vs. realism at the very top of the market.
Q: Can Cartesia Sonic clone a specific voice? A: Yes. Instant voice cloning from 10 seconds of audio is included on the $5 Pro tier and up. Professional-grade cloning (higher fidelity, longer intake) requires the $49 Startup tier or higher. Voice clones can be localized into 42 languages.
Q: How many languages does Cartesia Sonic support? A: Sonic 3.5 supports roughly 40 languages natively, with voice cloning localization available in 42. That covers all major commercial languages but trails ElevenLabs eleven_v3 (70+) and Scribe (90+) if you need coverage in less common markets.
Q: What does the Voice Agent add-on cost? A: Voice Agent calls are $0.06 per minute on every tier, with an additional $0.014 per minute if you use a Cartesia-provided phone number. This bills on top of your TTS credits — a busy call-center bot pays for both the synthesized audio and the agent runtime.
Q: Can Cartesia be self-hosted or run in a private cloud? A: Yes, on the Enterprise tier. Cartesia supports on-premise, VPC, and in-region deployment for compliance with HIPAA, SOC 2 Type 2, GDPR, and PCI. Self-serve tiers are cloud-only through Cartesia's infrastructure.
Q: Is Cartesia Sonic suitable for audiobook narration? A: It works, but ElevenLabs eleven_v3 currently produces better long-form emotive narration for audiobooks and video voiceover. Sonic's design priority is real-time conversational agents, where its 90ms latency matters — for a two-hour audiobook, that latency advantage is irrelevant.
Q: How is Cartesia billing "credits" different from ElevenLabs "characters"? A: ElevenLabs bills per character consumed, which makes cost forecasting straightforward. Cartesia bills in credits that translate to characters based on model and options used, which is more opaque. Expect to run a representative workload before you can confidently forecast monthly spend.