NEWS · SEPTEMBER 28, 2026 · CREATIVE

ElevenLabs launches Eleven v4: 90+ languages, voice clones from 10 seconds

On September 28, 2026 ElevenLabs announced Eleven v4, a new text-to-speech model, and Eleven v4 Turbo, a variant built for latency-sensitive work such as agents. The company calls v4 its most emotive model yet. Both support more than 90 languages, Instant Voice Clones can now be built from 10 seconds of audio, and Turkish is on the documented language list.

01 · WHAT HAPPENED?

A new architecture, two models and three ways in

The announcement went up on the ElevenLabs blog on September 28, 2026 under the names of Mati Staniszewski and Piotr Dabkowski, with an update on September 29. It introduces two models at once: Eleven v4 and its low-latency sibling, Eleven v4 Turbo. Eleven v4 sits on an entirely new architecture and was designed to read tone, pacing, emotion, character and context from the text it is given. As the post puts it, "Both models are available now in ElevenAgents, ElevenCreative, and via ElevenAPI."

Webrazzi covered the launch in Turkish on September 29, 2026, in a piece by Tuğçe İçözü. It reports that the previous Eleven v3 supported about 70 languages and that ElevenLabs claims significant quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese with the new release. The company's model documentation tells the same story in table form: eleven_v4 and eleven_v4_turbo are listed at 90+ languages, eleven_v3 at 70+.

02 · DETAILS

The script directs the delivery, and voices hold up over long jobs

The most visible improvement is in direction. Users can describe in plain language how a line should be read and drop inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing] into the text to add emotion, accent or sound effects. ElevenLabs says v4 follows these tags and direction prompts more accurately than earlier models. Webrazzi notes that inline tags arrived with Eleven v3 and that v4 extends the system, adding that several tags can be combined in one prompt and applied in the order given. The company also says it has markedly improved how the model handles International Phonetic Alphabet (IPA) phonemes, making custom pronunciations more dependable. The documentation caps a single eleven_v4 request at 10,000 characters, roughly 10 minutes of audio, against 5,000 characters for eleven_v3. Request stitching, which chains generations together for longer content, is also more reliable, and the company says that improves both ElevenLabs Studio and the ElevenLabs Reader app.

On cloning, Instant Voice Clones can now "capture voices with high fidelity using just 10 seconds of audio", according to the post, and v4 adds support for Professional Voice Clones for the highest-fidelity work. Webrazzi notes that Professional Voice Clones were not supported in Eleven v3 and are returning with the new model. According to the company, voices keep their identity more consistently through dialogue, narration and re-generated lines, and a voice recorded in one language can speak others with a native speaker's accent while keeping its original identity. Eleven v4 Turbo was optimised together with ElevenAgents, the company's conversational agents platform, as a single system. ElevenLabs gives Turbo a median inference latency of about 100 ms and a median time to first speech of about 150 ms; the footnote on the second figure says it was measured in September 2026 with identical scripts and default settings, network latency removed, and Turbo running over WebSocket streaming. Webrazzi reports that bidirectional streaming lets Turbo start producing audio before a sentence is fully generated, and that voice generation can begin before the background LLM response is complete.

03 · WHY IT MATTERS

Delivery becomes something you write, while the performance claims are the vendor's own

ElevenLabs lists audiobooks, character performances, voiceovers, dubbing and localising conversational agents as the target uses. For a business producing ad films, product videos or customer service agents, the practical upshot is concrete: tone can be steered with tags and descriptions inside the script, a character voice stays steadier across a long production, and a chosen brand voice can be carried into other languages. The flip side is that a clone from 10 seconds of audio makes the question of whose voice is used, and with what permission, more pressing on every project.

The performance claims come from the company itself. ElevenLabs says v4 ranks first on the Artificial Analysis Provider Voice Arena leaderboard for September 2026, and that about 75 percent of listeners preferred it in blind head-to-head tests against Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite TTS and Google Gemini 3.8 Flash TTS. Webrazzi also points out that the latency figures rest on tests the company ran.

04 · TÜRKİYE

Turkish is supported, but there is no Turkish-specific quality claim

The sources contain two facts about Türkiye. First, the ElevenLabs model documentation lists Turkish (tur) among the 90+ languages supported by the Eleven v4 family; the same documentation also lists Turkish for the previous Eleven v3, so support itself is not new with v4. Second, the Türkiye-based technology outlet Webrazzi covered the launch in Turkish on September 29, 2026. Being on the list is not a quality statement in itself. ElevenLabs makes its quality claims across languages in general: it says v4 handles rhythm, emotion and delivery better in the languages it covers and that a chosen brand voice will sound good in any language v4 supports, but the announcement does not single out Turkish. The languages Webrazzi names for quality gains are Japanese, Brazilian Portuguese, Mandarin and Cantonese. The announcement does not mention any Türkiye-specific availability terms or restrictions either.

What follows is UNALSOFT commentary. How well Turkish voiceover really works, especially on brand names, proper nouns and abbreviations, only shows up when you test with your own script, and the improved IPA support gives you a tool for handling those pronunciation problems. We covered another move in Turkish speech technology, the Turkish-focused Alania and Duyu models from the startup Patientdesk.ai, on September 25; that article is independent of this announcement. With the cloning threshold down to 10 seconds, written and explicit consent for voice use in ad and content work matters even more. That conclusion comes from our own assessment, not from the sources.

UNALSOFT's take

We read this announcement through the lens of ad film and product video production. Being able to describe delivery inside the script, keep one character voice constant across every piece of a campaign and carry a chosen brand voice into other languages directly changes the workflow on jobs that turn out many variations quickly. That is why on the AI ad films side we treat listening to and approving the Turkish voiceover against the project's own script, and getting explicit consent from the voice owner, as separate steps. Because the performance and latency figures are the vendor's own claims, we recommend choosing a tool only after hearing the same script in several of them.

Which language and which tone should your ad film speak in?

A short conversation is enough to plan the tool, the consent process and a Turkish pronunciation test together for any campaign that needs voiceover, dubbing or character voices.

Message on WhatsApp