Codex MCP tutorial · generate

How to generate speech audio from text with Codex MCP and Medux

Use Codex and Medux through MCP to generate speech from text with a reusable voice asset. This guide covers input preparation, the approval gate, task monitoring, and a concrete output check.

POST/api/app/v1/generate/audio

Generate speech audio from text using a reusable voice asset.

Open the Create Audio API reference →

Codex-specific workflow

Run Generate Speech from Text as a Codex MCP task

Use the active workspace as the source of truth. Ask Codex to identify the text, voice asset ID, language, and supported speech settings, select medux_speech_create_audio, and keep the Medux task ID and returned output with the rest of the project. Naming the input roles and the expected result prevents an agent from guessing which file should be used where.

Codex prompt pattern

In this workspace, use medux_speech_create_audio through Medux MCP to generate speech audio from text.
First inspect and identify the text, voice asset ID, language, and supported speech settings.
Show me the exact tool arguments before execution.
After the task completes, return the generated speech audio and verify that the audio speaks the complete text with the intended reusable voice.

Approval gate

Before approving the call, compare the selected tool, file or asset IDs, ordering, timing, and output settings with the request. If a local filename was mapped to an uploaded Medux file ID, keep that mapping visible so it can be audited later.

Result verification

Let Codex monitor an asynchronous task until it reaches a terminal state, then save or report the generated speech audio. The final check is explicit: confirm that the audio speaks the complete text with the intended reusable voice.

What this tutorial does

Medux operation

Generate speech audio from text using a reusable voice asset.

Agent workflow

Codex selects medux_speech_create_audio, presents the arguments for review, calls Medux, and reports the response.

Before you start

Authentication

  • Authentication is required via Authorization header.

Preconditions

  • Requires a voice asset id.
  • Requires voice asset id from your Media Assets or official public assets, plus non-empty text.
  • Also accepts optional links and model.

Prompt Codex

Start with a clear instruction that names the Medux operation and the intended inputs.

Generate speech audio from this text using voice asset id {voice asset id} and model id speech-standard.

Request fields to review

Confirm these values before approving the MCP tool call. The linked API page remains the source of truth for current limits and response semantics.

FieldTypeRequirementPurpose
linksarray<string>OptionalOptional reference links submitted from the Playground form.
modelstringOptionalGeneration model label or version selected in the Playground form.
model_idstringOptionalStable public Medux model identifier.
textstringOptionalProvide the value described in the API reference.
titlestringOptionalProvide the value described in the API reference.
voice_asset_idstringRequiredReusable voice asset used for text-to-speech generation.

Step-by-step workflow

01

Connect Medux MCP to Codex

Add the Medux remote MCP endpoint to Codex, provide an authorized credential, and confirm that the medux_speech_create_audio tool is available.

02

Prepare the required inputs

Prepare the source files, asset IDs, text, or settings required by the operation. Upload local media first when the tool requests file IDs.

03

Ask Codex to run Create Audio

Use one direct instruction with the intended values. Codex should select the Medux tool and show the proposed arguments before execution.

04

Review and approve the tool call

Check that the tool is medux_speech_create_audio and that each file ID, task ID, option, and output setting matches your request before approving it.

05

Check the Medux response

If Medux returns an asynchronous task, let the agent monitor it until completion; otherwise review the immediate response and save any result identifiers or output URLs.

Workflow recap

User goal
  ↓
Codex selects medux_speech_create_audio
  ↓
Review inputs and approve the tool call
  ↓
Medux runs Create Audio
  ↓
Codex reports the response and output

Frequently asked questions

Can Codex use the Medux Create Audio API through MCP?

Codex can call the medux_speech_create_audio Medux MCP tool to generate speech audio from text using a reusable voice asset. This guide covers the required inputs, approval flow, result handling, and the matching API reference.

Which MCP tool does this tutorial use?

This workflow uses medux_speech_create_audio. Review its arguments before approval and consult the API reference for the current contract.

Quick answer

How do I run Generate Speech from Text with Codex MCP?

Codex can use medux_speech_create_audio through Medux MCP to generate speech from text with a reusable voice asset. The workflow identifies the exact inputs, reviews the tool arguments, monitors processing, and verifies the returned result.

What does this Medux tutorial cover?

Codex can use medux_speech_create_audio through Medux MCP to generate speech from text with a reusable voice asset. The workflow identifies the exact inputs, reviews the tool arguments, monitors processing, and verifies the returned result.

Related topics: Generate Speech from Text · Generate Speech from Text with Codex MCP · Medux MCP tutorial