Unified Codex + Claude tutorial · generate

Create an AI lip sync video with Codex or Claude MCP

Create an AI lip sync video through the Medux API with Codex or Claude MCP. Upload avatar video and driving audio, review the call, and verify mouth sync.

Watch the complete workflow: create an audio-driven avatar video with Codex, Claude, MCP, and Medux. Watch on YouTube ↗
Choose your AI client

Follow the Codex or Claude workflow

The Medux operation and request fields are shared. Switch tabs for the client-specific prompt, approval pattern, screenshots, and walkthrough.

Codex-specific workflow

Run AI Lip Sync Video as a Codex MCP task

Use the active workspace as the source of truth. Ask Codex to identify the source avatar video and driving audio, with each role clearly assigned, select the Medux lip-sync tool, and keep the Medux task ID and returned output with the rest of the project. Naming the input roles and the expected result prevents an agent from guessing which file should be used where.

Codex prompt pattern

In this workspace, use the Medux lip-sync tool through Medux MCP to create an AI lip sync video.
First inspect and identify the source avatar video and driving audio, with each role clearly assigned.
Show me the exact tool arguments before execution.
After the task completes, return the lip-synced MP4 and verify that mouth movement follows the driving speech and the full video and audio play correctly.

Approval gate

Before approving the call, compare the selected tool, file or asset IDs, ordering, timing, and output settings with the request. If a local filename was mapped to an uploaded Medux file ID, keep that mapping visible so it can be audited later.

Result verification

Let Codex monitor an asynchronous task until it reaches a terminal state, then save or report the lip-synced MP4. The final check is explicit: confirm that mouth movement follows the driving speech and the full video and audio play correctly.

Before you start

Keep the setup small. The operation video uses a short avatar clip and a WAV audio file, then asks Codex to operate Medux through the MCP connection.

1. A source avatar videoUse a clear front-facing or slightly angled video. The face should be visible, stable, and not heavily blocked by hands, captions, or UI overlays.
2. A speech audio fileUse a clean WAV or MP3 file. The audio drives the mouth timing, so avoid long silence, heavy background music, or overlapping voices.
3. Codex with MCP accessConnect Codex to the Medux MCP server using the configuration from your Medux account or internal setup guide.
4. A clear output nameSave the result with a simple name such as lip_synced_avatar.mp4 so it is easy to preview and reuse.
# Example project layout
project/
  avatar.mp4
  reference_voice.wav
  tts.txt
  config/
  output/

The Medux MCP flow

The operation is straightforward: Codex reads your goal, checks the files, chooses the Medux tools, uploads media, creates the generation task, polls the task status, and saves the final video locally.

Codex reads the request
MCP connects to Medux
Files are uploaded
Medux renders
MP4 is returned

Step-by-step tutorial

01

Open the project in Codex

Start inside the folder that contains your avatar video and audio file. In the operation video, Codex first checks the workspace so it knows exactly which files can be used for the Medux task.

Codex project screen with the lip-sync request
Codex receives the lip-sync request and prepares to inspect the local assets.
02

Connect Medux as an MCP tool

Make sure the Medux MCP server is enabled before asking Codex to run the workflow. Once MCP is connected, Codex can call Medux directly instead of making you switch dashboards or write a separate upload script.

Tip: Keep your API key outside the prompt. Store it in your local MCP configuration or environment variables.
# Example only: use the exact MCP config from your Medux account
export MEDUX_API_KEY="your_medux_api_key"

# Then start Codex in the folder that contains your media files
codex
03

Ask Codex to run the lip-sync job

Use one direct instruction. The best prompt names the input files, explains the goal, and tells Codex what to return. This keeps the tool call clean and reduces back-and-forth.

Use Medux through MCP to create a lip-synced avatar video.

Inputs:
- avatar.mp4
- reference_voice.wav

Please upload both files to Medux, create the required avatar source, generate a lip-and-tongue synced video from the audio, poll the task until it is complete, and download the final MP4 as output/lip_synced_avatar.mp4.
Codex checking files for the Medux lip-sync workflow
Codex reviews the workspace and prepares the assets for the Medux request.
04

Let Codex upload the media and create the source avatar

Codex uses the Medux MCP tools to create upload slots, send the video and audio, and prepare the avatar source. In practice, this is the part that replaces manual dashboard uploads.

Codex using Medux MCP tools to upload media
The video shows Codex moving through the upload and preparation phase.
05

Run the audio-driven talking-head task

After the avatar source is ready, Codex launches the generation task. Medux uses the video as the visual identity and the audio as the timing reference, then renders a synced output.

Codex creating the Medux lip-sync generation task
Codex starts the Medux generation task and keeps the workflow inside one interface.
06

Poll the task until Medux returns the result

Do not refresh manually. Codex can keep checking the Medux task status and continue only when the render is finished. This is especially useful for longer videos or higher-quality outputs.

Codex polling the Medux task status
Codex tracks the task status until Medux finishes rendering.
07

Download and preview the final MP4

When the task is complete, Codex downloads the output and saves it locally. Open the result and check three things: the mouth follows the audio, the face remains stable, and the final video plays from start to finish.

# Quick local check
ffprobe output/lip_synced_avatar.mp4

# Preview the generated video
open output/lip_synced_avatar.mp4
Preview of the final lip-synced avatar video
The final output is saved locally and can be previewed immediately.

Quality checklist

Use this checklist before publishing the result or adding it to a larger content workflow.

Face visibilityThe face should remain visible throughout the clip. Avoid shots where the mouth is covered or too small.
Audio claritySpeech should be clean and centered. Remove long silence before sending the audio to Medux.
Timing reviewWatch the first few seconds and a later section. Make sure the mouth still follows the voice after the midpoint.
Final exportUse a practical MP4 filename, then keep both the source files and the output in the same project folder for repeatability.

Publish-ready summary

Medux gives Codex a media workflow layer through MCP. For lip and tongue sync, Codex only needs the source video, the speech audio, and a clear instruction. Medux handles the rendering, while Codex handles the tool calls, task tracking, and final download.

No separate dashboard hopping
No manual API upload script
One prompt to final synced video

Example result

Final lip-sync result

Below is the final lip-synced avatar video generated by Medux after Codex completes the MCP workflow.

rst
Generated result. Final lip-sync video generated by Medux.

Claude-specific workflow

Run Create an Audio-Driven Avatar Video as a Claude MCP workflow

Keep the goal, source assets, and constraints together in the conversation. Ask Claude to restate the reusable avatar ID, driving audio, and supported video settings before proposing medux_video_create_audio_driven. That context checkpoint makes the file roles and desired outcome easy to correct before any Medux call is approved.

Claude prompt pattern

Using the assets and requirements in this conversation, help me create an audio-driven avatar video with Medux MCP.
Restate which input fulfills each role: the reusable avatar ID, driving audio, and supported video settings.
Propose the medux_video_create_audio_driven call and summarize its arguments before asking for approval.
When it finishes, return the generated avatar video with a checklist confirming that the avatar uses the intended audio and the returned video plays from start to finish.

Context checkpoint

Have Claude summarize the intended transformation, identify every source asset by role, and list the important constraints. Review that summary together with the proposed tool arguments; correct the conversation first if a file, order, time range, or setting is ambiguous.

Result verification

Keep the Medux task ID in the conversation while Claude checks progress. When processing ends, ask for the generated avatar video plus a concise validation checklist confirming that the avatar uses the intended audio and the returned video plays from start to finish.

What this tutorial does

Medux operation

Generate an avatar video by pairing a reusable avatar with an uploaded audio file.

Agent workflow

Claude selects medux_video_create_audio_driven, presents the arguments for review, calls Medux, and reports the response.

Before you start

Authentication

  • Authentication is required via Authorization header.

Preconditions

  • Prepare avatar and audio assets before calling this endpoint.
  • Requires avatar asset id from your Media Assets or official public assets, plus audio file id.
  • Also accepts optional duration seconds, links, clarity, aspect ratio, and model.

Prompt Claude

Start with a clear instruction that names the Medux operation and the intended inputs.

Create a 7 second 1080P 16:9 talking video using avatar asset id {avatar asset id} and audio file id {audio file id}. Use model id talking-video-audio-standard.

Request fields to review

Confirm these values before approving the MCP tool call. The linked API page remains the source of truth for current limits and response semantics.

FieldTypeRequirementPurpose
aspect_ratiostringOptionalRequested output aspect ratio, such as 16:9 or 9:16.
audio_file_idstringRequiredProvide the value described in the API reference.
avatar_asset_idstringRequiredReusable avatar asset ID owned by the current user or listed as an official public asset.
claritystringOptionalRequested output resolution label, such as 1080P or 2K.
duration_secondsint32OptionalRequested output duration, in seconds.
linksarray<string>OptionalOptional reference links submitted from the Playground form.
modelstringOptionalGeneration model label or version selected in the Playground form.
model_idstringOptionalStable public Medux model identifier.
titlestringOptionalProvide the value described in the API reference.

Step-by-step workflow

01

Connect Medux MCP to Claude

Add the Medux remote MCP endpoint to Claude, provide an authorized credential, and confirm that the medux_video_create_audio_driven tool is available.

02

Prepare the required inputs

Prepare the source files, asset IDs, text, or settings required by the operation. Upload local media first when the tool requests file IDs.

03

Ask Claude to run Create Audio Driven Video

Use one direct instruction with the intended values. Claude should select the Medux tool and show the proposed arguments before execution.

04

Review and approve the tool call

Check that the tool is medux_video_create_audio_driven and that each file ID, task ID, option, and output setting matches your request before approving it.

05

Check the Medux response

If Medux returns an asynchronous task, let the agent monitor it until completion; otherwise review the immediate response and save any result identifiers or output URLs.

Workflow recap

User goal
  ↓
Claude selects medux_video_create_audio_driven
  ↓
Review inputs and approve the tool call
  ↓
Medux runs Create Audio Driven Video
  ↓
Claude reports the response and output

Frequently asked questions

Can Claude use the Medux Create Audio Driven Video API through MCP?

Claude can call the medux_video_create_audio_driven Medux MCP tool to generate an avatar video by pairing a reusable avatar with an uploaded audio file. This guide covers the required inputs, approval flow, result handling, and the matching API reference.

Which MCP tool does this tutorial use?

This workflow uses medux_video_create_audio_driven. Review its arguments before approval and consult the API reference for the current contract.

Quick answer

How do I create an audio-driven avatar video with Medux MCP?

Connect either Codex or Claude to Medux MCP, prepare the required inputs, review the proposed medux_video_create_audio_driven call, and verify the returned result. The tabs above keep client-specific prompting on one canonical page.

For a broader production workflow around approved audio, timing, and mouth-motion review, see the AI lip sync video guide.

Can I use the same Medux tool in Codex and Claude?

Yes. The client interaction differs, but both use the same Medux MCP tool and API request contract shown in this tutorial.

AI lip sync for developers

Build an AI lip sync video workflow with Medux API and MCP

Use this tutorial as a practical AI lip sync API workflow: provide an avatar video and a driving audio track, let Codex or Claude call Medux through MCP, and verify that the generated mouth movement follows the speech. The same pipeline supports audio-to-video lip sync, AI video lip sync, avatar lip sync, talking avatars, and approved dubbing workflows.

For production use, keep speaker consent, source ownership, audio quality, face visibility, timing, and final playback checks in the workflow. The resulting lip sync video can support localized content, training videos, and approved digital presenters.

Can I use Medux as a lip sync API with Codex or Claude?

Yes. Both AI clients can call the Medux audio-driven video tool through MCP, submit the avatar video and audio, monitor the asynchronous task, and return the lip-synced result.

What inputs does an audio-to-video lip sync workflow need?

Prepare a source avatar video and an authorized driving audio track. Clear speech, a visible face, and a stable source clip make output review easier.

Responsible lip sync for developers

Build a consent-based AI lip sync API for developers

Use authorized avatar video and driving audio, keep consent and source-ownership records outside the prompt, review the Medux tool arguments, and attach the accepted result to the originating task. This creates an auditable AI lip sync workflow for Codex MCP or Claude MCP without implying that consent verification happens automatically.

A security-conscious production setup keeps API keys out of prompts, limits media access, defines retention, and requires human review before publishing a synthesized likeness or voice.

Does Medux automatically verify avatar or voice consent?

This tutorial does not document automatic consent verification. The integrating application should collect permission, restrict access, retain the authorization record, and use only approved source media.