ElevenLabs Avatars is a new ElevenCreative entry point for making pre-rendered talking-head videos. Announced on June 11, 2026, it places persistent visual identities, ElevenLabs speech, and third-party lip-sync models in one interface.
The product is less about a new foundation model than a tighter generation layer. A creator selects an avatar, writes or reuses a script, chooses a voice, generates speech, previews it, and renders synchronized video without exporting audio between products. That convenience is meaningful, but the result is still generated media that requires consent, review, and asset management.
What an ElevenLabs Avatar is
An Avatar is a reusable visual identity representing a person, fictional character, or animal. It can be created from several reference images or from a text description. ElevenLabs recommends three to five high-quality images from different angles when references are used.
Once created, the identity can have multiple Styles. A Style changes visual context—camera angle, outfit, background, accessories, or lighting—while aiming to retain the underlying identity. Styles can be prompted or guided by another image.
Each Avatar can have a default voice, but a creator may select another voice from the workspace. The documentation includes community voices, Voice Design results, Instant Voice Clones, and Professional Voice Clones where the account and rights permit them.
The identity and voice are separate choices. Pairing them in one interface does not prove that the depicted person authorized the selected voice, or that a voice owner authorized the depicted character. A workspace needs evidence for both.
The combined voice-and-video workflow
The generation path is deliberately staged:
- select an Avatar and Style;
- open the lip-sync creation flow;
- select a voice;
- enter a script or choose a previous TTS generation;
- generate and listen to the speech;
- optionally guide the visual performance;
- render the synchronized video.
ElevenLabs says it automatically selects an appropriate lip-sync model based on input and quality requirements, although a creator can change the model in the generation step. This means “ElevenLabs Avatars” is a product surface that coordinates speech and visual models rather than one fixed video checkpoint.
That distinction matters for repeatability. Different lip-sync backends, resolutions, or Styles can produce different gestures and identity fidelity. Record the selected configuration with every approved output instead of treating the Avatar name as complete provenance.
What integration improves
External audio handoffs introduce mismatches: a changed script can leave an old video, a new voice can be rendered against the wrong take, and sample rates or trimming can alter synchronization. Producing speech and lip sync in one environment reduces those opportunities.
The creator can approve the spoken take before video rendering. Because ElevenLabs owns the speech layer, the platform can pass the exact audio used into lip sync. This is a stronger contract than manually downloading a file, editing it, and hoping no timing changed.
Integration does not guarantee perfect mouth shapes. Names, numbers, acronyms, fast delivery, laughter, overlapping sounds, and non-human faces remain challenging. Check the entire output at normal speed and frame by frame around important phonemes.
Reproducibility needs more than an Avatar name
An Avatar is a reusable identity, not a frozen rendering recipe. The chosen Style, source images, voice, script, speech take, lip-sync backend, resolution, and product defaults all influence the result. A platform update may also change automatic model routing without changing the identity’s visible name.
For repeatable campaigns, create a versioned manifest for every release. Include the Avatar and Style identifiers, reference-set version, exact script, voice identifier, speech-generation identifier, lip-sync model when exposed, output settings, generation time, and checksums. Save a short approved reference clip for visual regression review.
Regenerate only with an explicit reason. If a correction affects one language or sentence, preserve the rejected asset and link it to the replacement. This makes it possible to explain why two videos using the same Avatar look or move differently.
Persistent identity is useful—and demanding
Reusable identities support recurring presenters, localized training, product explainers, personalized outreach, and character-led social content. A Style library can preserve an approved visual vocabulary across many scripts.
Persistence raises the cost of a mistake. If an Avatar drifts in one Style or an unauthorized reference enters the identity set, that problem can propagate into many videos. Maintain:
- the source images and their consent records;
- an owner and expiry date for the identity;
- approved and prohibited Styles;
- voice permissions and geographic restrictions;
- a list of allowed use cases;
- every published output and revocation path.
For a real person, define whether the Avatar may speak new words, endorse products, appear in sensitive contexts, or be localized into languages the person does not speak. Consent should be specific and revocable, not inferred from possession of photographs.
Availability, Flows, and API status
ElevenLabs documents Avatars on all paid plans. Some models and reference-image capabilities are restricted in the United States because of regulatory or provider requirements. Availability can therefore depend on region and selected backend.
The product supports an Avatar node in ElevenLabs Flows. That node can automate variations of scripts, voices, Styles, or campaigns inside the ElevenCreative automation environment.
Direct Avatar API access was not available at launch. ElevenLabs says it is planned for the future. Do not build an external application around an assumed endpoint or present Flows as equivalent to a public REST API. A future API may expose different models, limits, or consent requirements.
Credits and production cost
Avatar videos draw from the existing Image & Video credit system. Cost varies by lip-sync model, duration, and resolution, and all ElevenCreative products share an account’s credit pool. ElevenLabs does not publish one universal “price per Avatar minute” in the capability page.
Estimate a campaign using the actual configuration in the account. Track credits for speech retries separately from video renders. Approve pronunciation, pacing, and script before the more expensive visual stage, and avoid generating high-resolution versions of rejected takes.
Cost per accepted minute should include reference preparation, Style generation, voice work, lip-sync renders, review, captions, corrections, and storage. A one-tab experience can still contain several billable model calls.
Quality review should remain layered
Review audio before video:
- verify every word, name, number, and disclaimer;
- check pronunciation, accent, pace, emotion, and pauses;
- confirm the voice is authorized for the script and locale;
- listen for clipping, artifacts, and unintended sounds.
Then review the video:
- identity, skin, hair, clothing, and background stability;
- mouth closure, teeth, tongue, and jaw movement;
- gestures that match the delivery;
- eye contact, blinking, and framing;
- unwanted text, logos, or visual changes;
- duration, dimensions, codec, and audio sync.
Keep the exact approved script and audio checksum beside the final video. If a script changes, invalidate the associated render rather than updating a text record in place.
When separate generation layers are preferable
The integrated product is attractive when speed and a reusable identity matter. Separate tools remain valuable when a team needs a specific TTS model, a precise audio edit, a lip-sync engine with different controls, or an auditable boundary between voice and face.
Medux offers one such compositional path when called by an authorized Codex or Claude MCP client. It is not embedded in ElevenLabs Avatars. The Codex MCP TTS tutorial shows generating a speech asset as its own task, and the lip-sync tutorial shows using approved audio and visual inputs in a later task.
That separation lets a reviewer approve the voice file before any face is animated. Store signed-upload references, task IDs, source checksums, the approved script, and the downloaded result. If lip sync fails, retry that operation without silently regenerating the voice.
ElevenLabs Avatars optimizes for a coherent creative surface; a composed Medux workflow optimizes for explicit boundaries and replaceable stages. Neither approach removes the need for permission and review. The right choice depends on whether the organization values one integrated workspace or tighter control over every artifact handoff.