ElevenLabs has released Eleven v4 and Eleven v4 Turbo, a new generation of text-to-speech models that can start an Instant Voice Clone from as little as 10 seconds of audio. The company is also targeting more natural long-form speech, multilingual output and faster responses for real-time voice agents.

Eleven v4 is designed to better preserve a speaker’s vocal identity while accounting for tone, pacing, emotion, character and context. ElevenLabs says the model addresses common synthetic-speech shortcomings including uneven cadence, misplaced emotion, shifts in timbre and disconnected-sounding sentences in longer passages or multi-speaker dialogue.

The models also let creators place performance directions directly within the text, extending control beyond the words in a script.

  • Nonverbal cues such as [laughs] and [whispers]
  • Emotion and pacing instructions
  • Accent directions
  • Sound effects, including rain or doorbells

The 10-second figure applies to Instant Voice Clone, rather than ElevenLabs’ more involved Professional Voice Clone process, which is intended to prioritize maximum fidelity. The lower input requirement could make synthetic narration, game characters, AI assistants and content localization faster to produce, while also lowering the barrier to reproducing a voice.

Turbo targets live voice AI

Eleven v4 Turbo is aimed at voice agents and live interactions, where response delays can break the flow of a conversation. ElevenLabs cites roughly 100 ms median inference latency and around 150 ms to first audio for the Turbo model.

Both models support more than 90 languages. ElevenLabs says they can retain the source voice’s identity while adapting pronunciation to the selected language, a claim that will be particularly relevant for long-form multilingual speech rather than short demonstration clips.

The launch arrives as ElevenLabs’ technology is appearing inside larger creative and media products: Adobe has integrated ElevenLabs voices into Firefly’s Generate Speech feature, while Spotify is working with the company on AI-narrated audiobooks.

Eleven v4 and Eleven v4 Turbo are available through ElevenCreative, ElevenAgents and ElevenAPI. The practical test for the new models will be whether their claimed voice consistency and low latency hold up in extended, multilingual conversations and production workflows.

SOURCEelevenlabs.io
Previous articleMotorola’s rumored wide foldable may pair a 165Hz display with a 5,000mAh battery