Industry-standard text-to-speech and voice cloning. v3 adds emotion control, multi-character dialogue, and 32 languages with native accent quality.
Audio & Voice
ElevenLabs v3 is the company's flagship expressive text-to-speech model, moved from public alpha (June 2025) to general availability on March 14, 2026. It replaces Multilingual v2 as the default recommendation for anyone doing narration, character work, or long-form audio where the voice needs to actually sound like it means what it says. The headline capability is inline audio tags — bracketed cues like [laughs], [whispers], [excited], [sighs], [sarcastic] — that steer prosody without needing a separate emotion parameter or a re-recording pass. ElevenLabs reports a roughly 68% drop in complex-text errors versus v2 (numbers, abbreviations, mixed-language passages), and the majority of blind-test users preferred v3 over the alpha build.
If you are a solo creator producing YouTube voiceovers, a small studio dubbing shorts into ten languages, or a game developer scripting NPCs, v3 is the current best-in-class for output that doesn't sound like a smart speaker reading a receipt. It is expensive versus Cartesia and OpenAI, and it is deliberately not built for sub-second real-time — those are the two honest catches. For everything else in the "make this text sound like a human performance" category, nothing else on the market lands the same on emotion, breath, and micro-pause timing.
The competitive field in 2026 is crowded. Cartesia's Sonic 3.5 wins on latency (roughly 166-190 ms end-to-end in independent tests) and undercuts on price at scale, which is why voice-agent builders keep picking it. OpenAI's gpt-4o-mini-tts and tts-1-hd are cheaper per character and easier to bolt onto an existing OpenAI stack, but flatter in delivery. PlayHT and Resemble AI compete on voice cloning workflows for studios. Murf AI targets corporate e-learning with a stiffer, more "presenter" delivery. ElevenLabs v3 sits above all of them for expressive quality — which is exactly what a novelist narrating an audiobook or an ad agency cutting a spot is paying for.
Audio tag markup for prosody control. The core v3 innovation is a set of inline bracketed cues you drop into the script itself. [laughs], [whispers], [sighs], [excited], [curious], [sarcastic], [nervous] and a growing library of dialect and mood tags rewrite the delivery of the surrounding sentence rather than triggering a canned sound effect. Compared to v2's monolithic "stability / similarity / style" sliders, this is closer to writing a screenplay: the actor reads the parenthetical. It also means edits happen in the script, not in a dashboard, which matters when you're revising a 40-minute chapter.
Dialogue mode for multi-speaker scenes. v3 accepts multi-speaker input in a single generation and preserves distinct voices, turn-taking, and overlap timing. This replaces the earlier workflow of generating each character's line separately and stitching in a DAW. The trade-off: dialogue mode still occasionally bleeds prosody from one speaker into another on tight cuts, and ElevenLabs recommends re-rolling those seams rather than crossfading. It is genuinely useful for radio-drama-style content and game cutscenes.
70+ language support with cross-lingual voice cloning. A voice cloned from a 30-second English sample will speak Japanese, Arabic, Turkish, or Hindi in the same voice. Quality is uneven — English, Spanish, French, German, Portuguese, and Italian sound essentially indistinguishable from a native performance; Arabic, Hindi, and Mandarin are noticeably better than v2 but still tell on themselves in fast passages; low-resource languages like Malay or Amharic are usable but wooden. For dubbing pipelines this is the single biggest reason to pick ElevenLabs over Cartesia (which currently trails on non-English coverage).
Professional Voice Cloning vs Instant Voice Cloning. Instant Voice Cloning (Starter tier and up) needs about 30 seconds of clean audio and produces a serviceable clone in under a minute. Professional Voice Cloning (Creator tier and up) ingests 30 minutes to 3 hours of source audio, takes several hours to train, and produces a clone that holds up in long-form narration — audiobook publishers actually license these. The gap between the two is not marketing; a PVC clone at v3 quality is the current bar for AI narration passing casual listener detection.
Two-model strategy: v3 for expression, Flash v2.5 for real-time. ElevenLabs is upfront that v3 is not the model for a live voice agent — the larger network and higher-fidelity codec push generation latency well past the ~75 ms Flash v2.5 achieves. If you're building a phone bot or a real-time conversational avatar, ElevenLabs itself points you at Flash v2.5 or the ElevenLabs Agents product. That is the honest split; a lot of buyers show up expecting v3 quality at Cartesia latency and don't get it.
Studio and Projects for long-form work. Studio (the successor to Projects) is a browser-based editor for turning a manuscript into a stitched, chapter-marked, mastered audio file. It handles pronunciation dictionaries, per-paragraph voice assignments, retakes at the sentence level, and export to WAV / MP3 / chaptered M4B. It is the reason serious audiobook narrators shortlist ElevenLabs — no other TTS vendor ships a comparable long-form editor.
API with 44.1 kHz PCM and streaming. Pro tier ($99/mo) unlocks 44.1 kHz PCM output over the API and 192 kbps quality — the point below which serious broadcast and podcast workflows won't touch you. Below Pro you're capped at MP3 quality. Streaming, chunked generation, and websocket endpoints are supported for building your own real-time layer (with Flash v2.5, not v3).
Sound Effects, Music, and Dubbing as bundled products. Every paid tier now includes access to ElevenLabs' text-to-sound-effects generator, Music v2 (text-to-song), and Dubbing Studio (auto-translate + lip-sync-timed dubbing of existing video). None of these are best-in-class individually — Suno and Udio beat Music v2 for song generation, HeyGen beats Dubbing Studio for lip-synced dubbing — but bundling them into the same credit pool means small teams don't need three subscriptions.
ElevenLabs prices by monthly "credits" (1 credit ≈ 1 character on v3, more on lighter models). Every paid tier includes commercial license and access to v3, Multilingual v2, Flash v2.5, sound effects, music, and dubbing — the differentiators are credit volume, voice-cloning tier, audio quality caps, and workspace seats.
| Tier | Price/mo | Credits | Key unlocks |
|---|---|---|---|
| Free | $0 | 10,000 | v3 access, non-commercial only, watermark on cloned-voice output |
| Starter | $6 | 30,000 | Commercial license, Instant Voice Cloning, Dubbing Studio |
| Creator | $22 | 121,000 | Professional Voice Cloning, additional credit purchasing |
| Pro | $99 | 600,000 | 44.1 kHz PCM API output, 192 kbps quality |
| Scale | $299 | 1,800,000 | 3 workspace seats, 3 PVC slots, team collaboration |
| Business | $990 | 6,000,000 | Low-latency TTS at $0.05/min, 10 PVC slots, 10 seats |
| Enterprise | Custom | Custom | SSO, HIPAA BAA, DPA/SLA, dedicated concurrency, managed dubbing |
For a solo creator doing weekly YouTube voiceovers (roughly 20-30 minutes of finished audio a week), Creator at $22/mo is the honest floor — Starter's 30,000 credits burn out in about 40 minutes of v3 output. For a studio publishing audiobooks or dubbing a channel into multiple languages, Pro at $99/mo is where the API-quality unlock lives; below that you cannot pull broadcast-grade files. Cartesia's Startup plan is $49/mo for 1.25M credits — nearly 10x more headroom for less than half the money — and OpenAI's tts-1 is $15 per 1M characters flat with no subscription. If your work needs the v3 emotion layer, you pay the ElevenLabs tax; if it doesn't, you're overspending.
Pros
Cons
Cartesia (Sonic 3.5) — Free / $5 Pro / $49 Startup / $299 Scale, plus $0.06/min for voice agents. Pick Cartesia if you're building a voice agent, IVR system, or anything where sub-200 ms end-to-end latency is non-negotiable. Cartesia's Sonic 3.5 measured at roughly 166-190 ms in independent testing; ElevenLabs v3 cannot match that architecturally. Pick ElevenLabs v3 if your output goes into a finished piece of content where quality beats speed.
OpenAI TTS (tts-1-hd, gpt-4o-mini-tts) — $15/1M characters for tts-1, $30/1M for HD, roughly $0.015/min for gpt-4o-mini-tts. Pick OpenAI if you're already inside the OpenAI stack, want no-subscription pay-as-you-go, and can accept a flatter, more "assistant"-sounding delivery. Pick ElevenLabs v3 if you need character voices, emotional range, or a specific cloned voice — OpenAI cannot clone.
PlayHT — roughly $39/mo Creator / $99/mo Unlimited (verify current pricing at playht.com). Pick PlayHT for its long-standing conversational Agents product and its ultra-realistic English voices, particularly for U.S.-focused commercial voiceover. Pick ElevenLabs v3 for wider language coverage, better non-English cloning, and the Studio long-form editor.
Resemble AI — $0.01/sec pay-as-you-go, $29/mo Creator, $99/mo Pro, Enterprise custom. Pick Resemble for enterprise voice cloning with heavy compliance requirements (they hold Netflix, Paramount, and Deutsche Telekom as customers) and the strongest security posture in the field. Pick ElevenLabs v3 for out-of-the-box quality without a solutions-engineering engagement.
Murf AI — Free, $29/mo Creator, $99/mo Business, custom Enterprise. Pick Murf for corporate e-learning, training videos, and internal explainers where you want a polished "presenter" voice and a template-driven editor. Pick ElevenLabs v3 for anything requiring character, emotion, or a specific cloned voice — Murf's stock voices are its ceiling.
Buy ElevenLabs v3 at Creator ($22/mo) if you're a solo audiobook narrator, YouTube voiceover creator, indie podcaster, or ad-agency copywriter — the emotional expressiveness and Studio editor pay for themselves in avoided re-records within the first month. Upgrade to Pro ($99/mo) the moment you need 44.1 kHz API output for broadcast or long-form publishing. Skip ElevenLabs and go to Cartesia if you're building a real-time voice agent, IVR, or phone bot where latency below 200 ms is a hard requirement. Skip and go to OpenAI TTS if you're inside the OpenAI stack and want zero-subscription usage billing. Skip and go to Resemble AI if you're an enterprise with HIPAA, SSO, and DPA gating your procurement. For every other "make text sound like a person actually performed it" job in 2026, v3 is the pick.
Q: Is ElevenLabs v3 free? A: ElevenLabs v3 is accessible on the Free tier at $0/month with 10,000 credits (roughly 10-15 minutes of v3 output), but Free tier output is watermarked on cloned voices and cannot be used commercially. Any commercial use requires at minimum the Starter plan at $6/month.
Q: How does ElevenLabs v3 compare to Cartesia Sonic 3.5? A: ElevenLabs v3 wins clearly on expressive quality, emotional control via audio tags, and non-English language coverage. Cartesia Sonic 3.5 wins on latency (roughly 166-190 ms end-to-end vs v3's much slower generation) and on price per character (Cartesia's $49 Startup plan gives 1.25M credits vs ElevenLabs' 121K at $22). Pick v3 for finished content, Cartesia for real-time agents.
Q: Can I use ElevenLabs v3 for real-time voice agents? A: No — ElevenLabs itself recommends Flash v2.5 for real-time and conversational use cases, which runs at roughly 75 ms latency but with lower expressive quality than v3. v3 is optimized for expressiveness, not speed, and the underlying model is too large for sub-second first-audio-byte in production.
Q: How good is ElevenLabs v3 voice cloning? A: Instant Voice Cloning (Starter tier, ~30-second sample) produces usable clones in under a minute — good enough for social content and prototypes. Professional Voice Cloning (Creator tier, 30 minutes to 3 hours of source audio) produces clones that hold up in commercially published audiobooks and are the current bar for AI narration that passes casual listener detection.
Q: What languages does ElevenLabs v3 support? A: v3 supports 70+ languages with cross-lingual voice cloning. Quality is strong for English, Spanish, French, German, Portuguese, Italian, and Japanese; good but detectable for Arabic, Hindi, Mandarin, and Korean; usable but stiff for low-resource languages like Malay, Amharic, or Welsh.
Q: Does ElevenLabs v3 have an API? A: Yes — the API is available from the Starter tier ($6/month) upward, but 44.1 kHz PCM output and 192 kbps quality are gated behind the Pro tier ($99/month). Below Pro you're capped at MP3 quality, which is insufficient for broadcast, professional podcast, or audiobook workflows.
Q: Is ElevenLabs v3 worth the price versus OpenAI TTS?
A: If you need emotional expressiveness, character voices, or specific voice cloning, yes — OpenAI cannot clone and its delivery is noticeably flatter. If you need generic narration in English and are already paying for OpenAI, tts-1 at $15 per 1M characters is dramatically cheaper and good enough for utility voiceover, help-center audio, or draft narration.
Q: Can I use ElevenLabs v3 output commercially? A: Yes, from the Starter tier ($6/month) upward — every paid tier includes a commercial license for standard TTS output. Voice-cloned commercial use requires you to have rights to the source voice, and HIPAA-regulated use requires an Enterprise BAA, which is not available on lower tiers.