ElevenLabs says Dubbing v2 can preserve the original speaker’s tone, pacing, delivery, and emotional intent across more than 90 languages. The model has a better basis for that claim than a transcript-only system because it conditions directly on the source performance rather than reconstructing expression from words alone.

But “preserve emotion” is not a binary capability. Emotion is carried by pitch, rhythm, loudness, timbre, pauses, breath, word choice, culture, face, and scene context. A dub can match the source pitch contour and still feel wrong to a native listener because the translated phrase is unnatural or the intensity does not fit the target language.

The useful answer is therefore conditional: Dubbing v2 can retain more performance information, but a team must test whether the intended emotion survives for its actual languages, speakers, genres, and audiences. Current ElevenLabs documentation also labels the model Alpha, which reinforces the need for evaluation.

What information survives beyond a transcript?

A transcript normally captures lexical content and punctuation. It may indicate a question or exclamation, but not the exact acceleration into a phrase, the length of a hesitation, a restrained laugh, or the difference between sincere and sarcastic emphasis.

Dubbing v2 conditions on the original audio. ElevenLabs says this transfers intonation, tone, timing, pace, and energy into the target speech. A direct performance signal can help preserve whether a line is calm, urgent, playful, intimate, formal, or angry.

The translated sentence still changes. Word order, syllable count, stress patterns, and conventional expressions differ across languages. The system must decide which acoustic cues to preserve and which to adapt so the target sounds natural. Exact acoustic imitation can be less faithful to the communicative intent than a culturally appropriate reinterpretation.

This is why “same emotion” should mean the same intended audience response, not identical pitch values.

Emotion has several measurable dimensions

Evaluate at least six properties separately:

A single quality score hides tradeoffs. A dub may sound like the speaker but use unnatural phrasing. It may convey the right emotion while drifting from identity. It may be linguistically excellent but start too late for the facial reaction.

Ask reviewers to score each dimension and add time-coded comments. The team can then choose whether to revise translation, regenerate speech, adjust the edit, or accept a small difference.

Ninety-plus languages is a coverage claim

ElevenLabs supports more than 90 languages, including many widely used global languages. That is valuable reach. It does not establish equal performance for every direction between those languages.

Some target languages have more available speech data, pronunciation resources, and translation examples. Some have complex honorifics, gender agreement, tones, or dialect variation. A source voice may map naturally to one language and sound accented or unstable in another.

Language labels are also broad. Spanish for Mexico, Argentina, and Spain differs in vocabulary and performance expectations. Arabic, Chinese, English, and Portuguese each cover major regional varieties. Specify locale, audience, formality, and pronunciation rather than choosing only a language name.

Build a launch tier. Validate high-priority languages deeply, release them under native ownership, and expand only after the process works. Do not generate all supported languages and assume menu coverage equals publication readiness.

Translation can preserve or destroy the performance

An emotionally accurate source line can become flat because the translation is too literal. Humor, understatement, politeness, and emphasis often require adaptation. A phrase that fits the duration may not fit the character.

Give translators access to the scene, speaker intention, character history, and target audience. Maintain a terminology guide but allow natural spoken phrasing. For brand content, distinguish terms that are legally fixed from language that can be adapted.

ElevenLabs says its sync-aware translation aligns starts, stops, and pacing. This reduces timing work, but it creates pressure to fit meaning inside the source duration. Review whether the chosen wording sacrificed nuance. Sometimes the better solution is a slightly different edit or a longer shot.

Back-translation can find omissions, but it cannot judge natural performance. Use native listeners who understand the product or story.

Source genre changes the challenge

A clean presenter speaking to camera is easier than overlapping drama. Emotional preservation becomes harder with whispers, shouting, crying, laughter, singing, interruptions, sarcasm, code-switching, crowd noise, or rapidly cut dialogue.

Marketing videos introduce another risk: enthusiasm may become exaggerated and make a claim sound more certain than the approved source. Training and safety content requires clarity over theatrical matching. Comedy may need a rewritten beat rather than acoustic transfer.

Group scenes require speaker separation and identity continuity. ElevenLabs recommends up to nine unique speakers per file for best quality and says overlapping speakers can be isolated. Test whether interruptions and reactions still make conversational sense after translation.

Create a difficult-scene set before committing an entire catalog. Include the maximum speaker count, worst source audio, strongest emotion, shortest cut, fastest line, named entities, and each required locale.

Visual emotion matters too

Audiences interpret emotion from the face and body as much as the voice. A dubbed line may have the right acoustic intensity but conflict with a smile, eye movement, gesture, or mouth shape.

Watch every result with picture. Check whether emphasis lands on a gesture, laughter coincides with expression, pauses match reaction shots, and sentence endings do not continue after a cut. The target audio should not make a neutral face seem to deliver an extreme performance.

Traditional dubbing accepts some lip mismatch because viewers understand localization conventions. For close-up advertisements or digital presenters, the threshold may be tighter. Decide whether timing adaptation, editorial changes, or synthetic lip adjustment is appropriate based on rights and audience expectations.

A rigorous evaluation design

Use a balanced test matrix: languages, language directions, locales, speaker genders and ages, voice qualities, emotional categories, intensity levels, genres, audio conditions, and single versus overlapping speakers.

Do not let reviewers see which model produced each candidate when possible. Include the source, a human dub or current production baseline, and at least one simpler transcript-conditioned system. Randomize samples.

Native raters should answer specific questions rather than “Which sounds best?” Ask what emotion they perceived, how intense it was, whether the speech sounded natural, whether the voice fit the same person, and whether any word or pronunciation was wrong.

Track inter-rater disagreement. Emotion perception varies by person and culture. A system is risky when reviewers strongly disagree, even if its average score is acceptable.

Evaluate regressions whenever ElevenLabs changes the Alpha model or workflow. Preserve source clips, target translations, output versions, settings, and approval results.

Operational status changes expectations

Dubbing v2 is available in ElevenCreative and through ElevenProductions. The self-service product supports creators and marketing teams; ElevenProductions adds human translators, voice casting, and professional mixing for higher-stakes work.

Current documentation identifies Dubbing v2 as Alpha and says its API is not yet live. Do not describe it as a generally available API or build unattended automation around an unreleased endpoint. Product behavior and pricing can change while the model matures.

For high-risk media, the managed path may be appropriate because emotional transfer and linguistic correctness need human direction. For routine creator content, self-service can work when a native reviewer owns each target version.

Performance-preserving dubbing can make it appear that a person spoke words in a language they do not know. Obtain explicit rights for voice synthesis, translation, territories, channels, and duration. Give the speaker or authorized representative access to the translated meaning.

Do not intensify emotion in a way that changes the person’s position. A calm caveat should not become an excited endorsement. A concerned warning should not become neutral. Preserve disclosure and keep records linking target audio to approved source and translation.

Protect source recordings and voice data as sensitive identity material. Limit access, define retention, and avoid sending unnecessary takes to any service.

When pitch correction helps—and when it harms

Pitch contributes to speaker identity and emotion, but a translated language may naturally use a different contour. Applying correction simply because the target does not match the source numerically can make speech robotic or culturally unnatural.

First identify the defect. If identity drift affects the overall register, a modest adjustment may help. If emphasis is wrong, regenerate or redirect the performance. If a phoneme sounds unnatural, fix pronunciation or translation. Pitch processing cannot repair meaning.

Always compare processed and unprocessed versions with native listeners. Check for metallic artifacts, flattened dynamics, altered formants, and lost emotional peaks.

Separate Medux pitch and lip-sync operations

Medux provides external pitch-adjustment and lip-sync tasks that Claude or Codex can call after independent MCP setup. They are not part of ElevenLabs Dubbing v2 and no native transfer is claimed.

The Claude pitch-adjustment tutorial shows a focused audio operation. The Codex lip-sync tutorial applies supplied audio to authorized video through an asynchronous job.

A controlled workflow preserves the approved Dubbing v2 output, documents the specific failure, applies one measured transformation, retains the task ID, polls status, and compares the result against the original. Lip sync needs consent for the depicted person and frame-by-frame review for mouth, teeth, face boundaries, and identity drift.

Do not automatically process every language. Extra transformations add cost, artifacts, and data transfers. Use them only when a native reviewer and picture review identify a real benefit. Track ElevenLabs charges, agent usage, and Medux credits separately.

Dubbing v2 can carry more emotional evidence than a transcript-only pipeline, which is a substantial step. Whether it preserves emotion across a particular set of 90-plus languages is not answered by the feature list. It is answered by native listeners, representative scenes, controlled comparisons, and transparent approval.