Google’s Gemini 3.5 Live Translate and ElevenLabs Dubbing v2 share a compelling promise: translate speech without flattening the person who delivered it. Both aim to carry rhythm, emotion, and vocal character across languages instead of returning neutral text read by an unrelated voice.
They are nevertheless built for different clocks. Gemini Live Translate operates during a conversation. ElevenLabs Dubbing v2 processes recorded media for later publication. “Which preserves performance?” cannot be answered without asking whether the performance must remain live or become an editable artifact.
Start with the source event
Gemini 3.5 Live Translate, launched in preview in June 2026, is an audio-only mode of the Gemini Live API. It continuously listens to a stream and returns translated speech with low delay. Google lists more than 70 languages and over 2,000 language pairs.
ElevenLabs Dubbing v2, introduced in May 2026 and documented as Alpha, takes an uploaded file or supported media source. It detects speech, identifies speakers, separates background sound, translates dialogue, synthesizes voices, and assembles a localized output. ElevenLabs lists more than 90 languages.
The practical boundary is clear:
| Requirement | Gemini 3.5 Live Translate | ElevenLabs Dubbing v2 |
|---|---|---|
| Main job | Simultaneous spoken translation | Post-production media localization |
| Published status | Preview | Alpha |
| Input pattern | Continuous live audio | Recorded audio or video |
| Output priority | Low-delay translated conversation | Reviewable dubbed media |
| Languages | 70+ and 2,000+ pairs | 90+ |
| Tools in translation mode | Not supported | Production controls in ElevenCreative or Productions |
| Real-time use | Yes | No |
| API at publication | Gemini Live API preview | Dubbing v2 API not yet live |
Status labels matter. Preview and Alpha indicate that behavior, availability, limits, or interfaces may change. Neither should be described as a universally stable production dependency without a current review.
What preserving a performance actually means
A translated performance has several independent dimensions:
- semantic meaning and terminology;
- speaker identity and vocal character;
- emotion, emphasis, and conversational intent;
- tempo, pauses, and phrase length;
- pitch movement and prosody;
- turn timing and overlap;
- background sound and spatial context;
- visual mouth movement when video is involved.
No single metric captures all of them. A live system may preserve an excited interruption but simplify a technical phrase. An offline dub may deliver better terminology and timing after review while losing the spontaneity of a two-person overlap.
Evaluation should compare the translated audio with the source and with a human reference. Native speakers need to rate meaning, naturalness, tone, speaker similarity, pacing, pronunciation, and cultural appropriateness. Video reviewers also need to inspect visible timing rather than assuming audio quality guarantees visual sync.
Gemini optimizes for conversational continuity
Live Translate processes speech directly to translated speech. Google says it preserves intonation, pacing, and pitch, which helps a joke remain playful, a warning remain urgent, or a hesitant answer remain hesitant. Continuous operation is better suited to meetings, travel, support, live interviews, and conversations where waiting for an exported dub would defeat the purpose.
Translation mode is intentionally narrow. Google’s documentation says it is audio-only and does not support tools, function calling, web search, additional instructions, thinking controls, context caching, or batch mode. That prevents teams from treating it as a general Gemini agent with translation added. The model’s job in this mode is the live language bridge.
The restrictions influence product architecture. Terminology enforcement, identity checks, CRM lookup, safety workflow, and meeting actions need to live outside the translation session unless Google expands the interface. Do not hide consequential business logic inside assumptions about an unsupported tool path.
Google applies SynthID to generated audio. Teams should also disclose machine translation to participants, provide a text fallback where possible, and preserve a clear path to a human interpreter for medical, legal, financial, emergency, or other high-stakes exchanges.
Live translation has distinctive failure modes
Google’s documentation identifies several limitations. A voice can shift after a pause, gender presentation can be wrong, and multiple speakers may collapse into one translated voice. Accents, closely related languages, language switching, background audio, and pauses can also produce artifacts or incorrect detection.
These are not cosmetic details. If two negotiators receive one voice, attribution may become unsafe. If the system chooses the wrong related language, a fluent-sounding translation may conceal a meaning error. If participants switch languages mid-sentence, turn continuity can break.
Test with the acoustic reality of deployment: conference-room echo, mobile networks, overlapping speakers, code-switching, names, numbers, abbreviations, and emotionally charged speech. Give users a way to repeat, slow down, spell, view text, or request a human. Never let smooth synthesized audio become evidence that the translation is correct.
ElevenLabs optimizes for editorial control
Dubbing v2 starts from a completed performance and can spend more time organizing it. ElevenLabs says it conditions the translation on the original delivery, carries emotion and rhythm, preserves distinct speakers, and handles background music or ambience separately from dialogue.
The offline process supports revision. A producer can inspect speaker assignments, edit translated text, correct names, adjust timing, regenerate a line, compare languages, and approve an export. That workflow is valuable for films, creator videos, courses, ads, product explainers, and podcasts where the output will be replayed many times.
Dubbing v2 includes a voice-cloning strength control. ElevenLabs documents a default of seven: increasing it can make the result resemble the original speaker more closely but may carry accent or reduce naturalness, while decreasing it can improve naturalness at the cost of resemblance. The right setting depends on language pair, speaker, and rights agreement rather than a universal maximum.
The documentation recommends limiting complex projects to roughly nine speakers for best quality. Productions with crowds, frequent overlap, songs, whispering, archival audio, or heavy sound design should run a proof segment before committing a full catalog.
Timing is not the same as lip sync
ElevenLabs calls Dubbing v2 sync-aware because it considers the timing and pacing of source dialogue. This can reduce rushed or stretched speech and improve fit within an existing scene.
It does not guarantee frame-accurate mouth shapes. Languages differ in syllable count, phonemes, word order, and visible articulation. A sentence can end at the correct time while the lips form incompatible movements throughout the shot.
For a podcast or off-camera narrator, timing may be sufficient. For a close-up, presenter, avatar, or advertisement, inspect every visible line and use a dedicated lip-sync process when needed. Also review cuts, breaths, room tone, music ducking, and sound effects; localization quality lives in the complete mix.
Rights and consent travel with the voice
Performance preservation increases the importance of permission. A translated voice can sound like the original speaker in languages they never spoke and statements they never personally recorded. Obtain consent that covers cloning, translation, target languages, territories, channels, duration, and revision.
Separate permission to record from permission to synthesize. A public video is not automatically a voice-cloning license. Keep source provenance, speaker releases, script approvals, model and setting versions, and export checksums. Provide a takedown and correction process.
For live translation, disclose the system to everyone whose speech enters the stream and apply regional recording rules. For dubbing, ensure that rights cover both the source media and every localized distribution. Children, employees, contractors, actors, and public figures may have additional contractual or legal protections.
Cost and latency should match the business outcome
A live conversation is evaluated per session: connection success, first translated audio, interruption behavior, semantic accuracy, and minutes of low-quality fallback. An offline dub is evaluated per finished minute: preparation, corrections, regeneration, mixing, quality control, and distribution value.
Do not compare a streaming audio price with a dubbing project price as though they buy the same outcome. Include human review, retry rates, storage, rights administration, and downstream editing. A fast mistranslation during a safety-critical call is expensive even if inference is cheap; a perfect studio dub delivered after an event cannot solve a live conversation.
A controlled Medux finishing path
Some teams will use a leading translation system for the language transformation and a separate production stack for an approved asset. Medux can participate in that offline finishing path, but this is not a claim of native Gemini Live Translate or ElevenLabs Dubbing v2 integration.
A traceable workflow can look like this:
- retain only authorized source material and produce an approved translated script;
- generate or replace a specific narration through the Medux TTS workflow when separate TTS is appropriate;
- adjust pitch only after editorial approval and without using it to imitate an unconsenting person;
- run a distinct Medux lip-sync task for visually critical footage;
- store source, script, consent, model settings, task IDs, and final checksums.
Use signed uploads and least-privilege credentials, and avoid sending raw live conversations into an offline media service automatically. Redact sensitive material first. If lip sync fails, retry that step rather than silently retranslating or replacing the voice.
Gemini 3.5 Live Translate is the stronger conceptual fit when people must understand one another now. ElevenLabs Dubbing v2 is the stronger fit when an audience must receive a polished recording later. Performance is best preserved when the system’s clock, controls, and review process match the actual medium.