ElevenLabs Dubbing v2 is built around a simple observation: a transcript does not contain a performance. It captures words, but usually loses timing, hesitation, emphasis, energy, breath, and emotional intent. Dubbing v2 conditions directly on the source performance so the translated voice can follow more than the text.
ElevenLabs launched the model on May 28, 2026 for more than 90 languages. It is available in ElevenCreative for self-service localization and through ElevenProductions for managed professional work. The company’s current documentation labels Dubbing v2 Alpha, and says the Dubbing v2 API is not yet live.
The title promise—multilingual video without losing performance—should be treated as the product’s objective, not a guarantee. Different languages reshape sentence length and rhythm, and visual speech creates constraints that audio alone cannot solve. A professional workflow still needs translation, native listening, identity and rights review, mixing, and picture approval.
Conditioning on performance instead of rebuilding from text
Traditional automated dubbing often follows three separate abstractions: transcribe the source, translate the words, then synthesize the translation. Once the first step turns speech into text, much of the original delivery is gone.
Dubbing v2 conditions on the source audio to capture intonation and tone. ElevenLabs says it preserves tone, pacing, delivery, and emotional intent across languages. The target is not merely a recognizable cloned voice reading a translation; it is a translated performance that behaves like the source speaker.
This matters for content in which delivery carries meaning. A quiet product explanation, an excited creator reaction, a comic pause, or a serious documentary statement can all become wrong when synthesized at a generic pace and emotion.
Performance conditioning does not remove the translator’s role. An idiom may need a different phrase, a joke may need replacement, and a formal register may not transfer directly. The dubbed performance must be faithful to the intended meaning, not just the acoustic contour.
Sync-aware translation handles language length
Languages do not express the same idea with the same number of syllables. A literal translation can be much longer than the source segment, forcing unnaturally fast speech or spilling across a cut.
ElevenLabs says its sync-aware translation adapts phrasing for spoken delivery and aligns starts, stops, and pacing. This allows the translation to breathe within the source timing rather than treating timing as a later stretch operation.
Review the compromise. A shorter translation may omit nuance. A longer one may rush. A phrase that matches duration may sound unnatural in a particular country. Give native reviewers the source meaning, intended audience, terminology list, and access to the picture—not just an isolated audio file.
Keep segment boundaries editable. If a sentence is repeatedly rushed, adjust the translation, move a cut, or allow a longer pause rather than applying extreme time compression.
More than 90 languages is coverage, not uniform quality
ElevenLabs advertises more than 90 languages. That breadth can make one campaign or channel accessible to audiences that could not justify a full traditional dubbing pipeline.
Coverage does not imply that every language pair, dialect, named entity, emotion, and speaker has equal quality. Source-to-target translation may behave differently when one language is low-resource or structurally distant. Code-switching, regional slang, technical terms, acronyms, numbers, and names need special attention.
Create a glossary with approved product names, pronunciations, units, and terms that should remain untranslated. Specify locale rather than only language: Brazilian Portuguese is not interchangeable with European Portuguese, and regional Spanish choices affect audience trust.
Test the hardest content first. If the project includes overlapping speakers, laughter, whispering, singing, poor microphones, or strong background music, do not base a schedule on one clean monologue demo.
Speaker separation and identity
Current ElevenLabs documentation says all audio and video content types are supported and recommends up to nine unique speakers per file for best quality. It also says overlapping speakers can be separated into tracks.
Multi-speaker content needs diarization checks. Confirm that each translated line belongs to the correct person, voices do not swap after a cut, and interruptions retain their conversational logic. Overlap that sounds natural in the source may become unintelligible in translation.
Voice identity requires authorization. Obtain consent or contractual rights to synthesize each speaker in the intended languages, territories, and media. A performer who agreed to one recorded language may not have agreed to a synthetic multilingual version.
Do not use a preserved voice to create new claims that the speaker never approved. The more convincing the performance transfer, the greater the risk of an audience assuming the person actually recorded the target speech.
Source preparation determines output quality
Start from the best available master. Preserve clear dialogue, stable sample rate, and enough separation from music and effects. Compression artifacts, clipping, reverberation, and background voices make transcription, speaker isolation, and performance analysis harder.
Keep music-and-effects stems when available. Professional localization often replaces dialogue while preserving ambience, score, and sound design. Trying to reconstruct the background from a flattened mix can introduce pumping or missing effects.
Provide a verified transcript. Even if the system creates one automatically, correct names, numbers, technical terms, and punctuation before translation. A beautifully performed mistranscription is still wrong.
Segment around meaning and edits. Avoid cutting a breath or sentence unnaturally simply to match an arbitrary subtitle boundary.
ElevenCreative and ElevenProductions serve different needs
ElevenCreative offers a self-service flow intended for creators and marketing teams. It can localize videos without coordinating a separate vendor for every stage. This is useful for channel content, tutorials, product videos, and early market tests.
ElevenProductions combines Dubbing v2 with human translators, voice casting, and professional audio mixing. Studios, broadcasters, and regulated brands may prefer this managed path when language accuracy, performance direction, rights, and final mix need specialist accountability.
The distinction is not simply budget. A short informal creator video may succeed with native review and self-service tools. A drama, major advertisement, safety instruction, or legal communication usually needs deeper human intervention.
Current documentation calls Dubbing v2 Alpha. Plan for occasional rough edges, keep source masters and rollback options, and avoid removing human QA because the user interface feels simple.
API status and automation limits
ElevenLabs’ launch post and current documentation say Dubbing v2 API access is coming soon. The product is not a live self-serve Dubbing v2 API at the time of this article. Do not build a production dependency around an endpoint that has not been released.
The existing ElevenLabs API catalog and legacy dubbing interfaces should not be assumed to expose the same v2 model or behavior. Verify model name, version, supported operations, pricing, concurrency, webhook behavior, and migration guidance when the API launches.
For now, automation should respect the available product surface. A human can prepare inputs, configure the job in ElevenCreative, review outputs, and export approved tracks. Enterprise teams can discuss managed delivery through ElevenProductions.
Quality assurance for a dub
Use at least four review passes:
- linguistic: meaning, terminology, grammar, locale, and omissions;
- performance: tone, emphasis, pace, pauses, emotion, and speaker identity;
- picture sync: starts, stops, mouth movement, cuts, gestures, and on-screen text;
- technical: clipping, noise, loudness, channels, sample rate, and final mix.
Native reviewers should watch the full video. A transcript-only review misses irony, facial expression, and actions that disambiguate language. Record time-coded issues and version every approved target language.
Do not force every voice to match the source pitch exactly. Natural pitch ranges differ, and aggressive correction can create artifacts or flatten expression. Evaluate whether a mismatch is an identity problem, a translation-performance problem, or simply a normal target-language variation.
Cost and operational planning
Dubbing cost depends on current plan or managed-services terms, duration, languages, and revision needs. The seven-day promotional allowances described at launch have expired by July 2026 and should not be used in forecasts.
Budget by approved target minute. Include source cleanup, transcription correction, translation, model use, native review, retries, mixing, lip-sync work, caption updates, storage, and distribution. A cheap initial dub that needs extensive repair can cost more than a managed workflow.
The documentation lists concurrency limits for self-service plans and advises waiting when too many jobs are active. When API access arrives, queues, idempotency, status tracking, and refund behavior will need explicit implementation.
Separate Medux finishing where needed
An approved Dubbing v2 track may still need pitch adjustment, visual lip synchronization, or attachment to a delivery video. Medux provides external processing tasks that Claude or Codex can call after independent MCP setup. It is not built into ElevenLabs and does not receive a dub automatically.
The Claude pitch-adjustment tutorial, Codex lip-sync tutorial, and Claude audio-track tutorial cover distinct asynchronous operations.
Use them selectively. Preserve the approved original dub, apply modest pitch changes only after listening tests, and perform lip sync only with authorization for the depicted person. Attach the final licensed track to the correct picture master and verify duration, sync, streams, and loudness. Retain each task ID and poll status rather than resubmitting a delayed job.
Keep ElevenLabs usage, agent tokens, Medux credits, and human review as separate costs. A multi-service chain also creates multiple data paths, so upload only approved media under current retention and privacy rules.
Dubbing v2 moves automated localization closer to performance translation rather than transcript replacement. Its real quality will be measured not by how many languages appear in a menu, but by whether native audiences hear the right meaning, person, and emotion at the right moment.