ElevenLabs launches Eleven v4 — its most emotive voice model yet, plus a Turbo for real-time agents
Published September 30, 2026. ElevenLabs has released Eleven v4, a text-to-speech model built on an entirely new architecture, alongside Eleven v4 Turbo, a low-latency variant aimed at voice agents and real-time use. Both are available now in ElevenAgents, ElevenCreative, and via API — including on the free tier.
What’s new
- A new architecture designed to interpret tone, pacing, emotion, character, and context — generating speech that can sound dramatic, tender, urgent, comedic, or conversational while holding the speaker’s identity.
- Inline audio tags replace SSML. You direct delivery in plain language — tags like [laughs], [said angrily in French accent], even [light rain] or [phone buzzing] — and the model follows them more accurately than its predecessors. Classic SSML break tags are disabled in v4.
- Better speaker-identity capture keeps voices consistent across agent conversations, audiobooks, and ads; scene-level context makes multi-line dialogue sound natural rather than stitched together.
- Improved IPA phoneme support for reliable custom pronunciations.
Why it matters
Voice is where “good enough” AI audio becomes genuinely usable — for faceless YouTube channels, audiobooks, and voice agents, the difference between flat narration and directed performance is the whole product. We already track ElevenLabs closely (see our ElevenLabs alternatives roundup and faceless YouTube guide), and v4 looks like the release that raises the bar for everyone else. Hands-on testing coming soon.
Source: Unite.AI.
