AI video used to arrive as a silent moving image. Sound effects, ambience, music, and dialogue were separate production problems. Gemini Omni Flash, Grok Imagine Video 1.5, and Meta Muse Video point toward a different default: the model imagines the audible scene at the same time as the visible one.
Native audio makes a concept easier to judge. Footsteps can land on movement, a door can sound when it closes, and spoken timing can shape the shot. It also combines several failure modes into one output. A convincing image can hide incorrect speech, an unlicensed musical idea, or synchronization that drifts by a few frames.
The three products are at different stages. Omni is a public-preview API. Grok 1.5 is generally available in the xAI API. Muse Video was previewed by Meta on July 7 and is described as coming soon, not as a service creators or developers can already depend on.
The current native-audio picture
| Model | Availability | Video contract | Audio position | Important boundary |
|---|---|---|---|---|
| Gemini Omni Flash | Public preview API | Text, image, references, and supported short-video tasks; 3–10 seconds at 720p | Attempts dialogue, ambience, and effects with video | No audio reference or voice editing |
| Grok Imagine Video 1.5 | GA API | Starting image, prompt, resolution, and duration; 480p–1080p | Effects, ambience, and dialogue in the same pass | Current API model is image-to-video |
| Meta Muse Video | Early preview, coming soon | Product specification not yet published | Meta says native audio is supported | Meta identifies audio-video sync as a performance gap |
Do not fill missing Muse specifications with assumptions. A preview sample is not an API contract, price, duration, or launch date.
What “native” does and does not mean
Native audio means the video-generation process produces an audio track as part of the same result. The model can use the prompt and visual trajectory to coordinate events. It does not mean the model recorded production sound, licensed a music cue, cloned an approved performer, or mixed to a broadcast standard.
Separate the output into layers during review even if the model returned one mixed track. Ask whether speech says the intended words, whether effects match visible causes, whether ambience fits location and perspective, and whether music is appropriate. A single pass can succeed on one layer and fail on another.
Native generation can also make revision expensive. If the picture is approved but one spoken word is wrong, regenerating the whole clip may change motion and composition. A production workflow needs a way to lock the visual and replace sound downstream.
Gemini Omni Flash: flexible input, limited voice control
Google’s Omni model can start from text, an image, tagged visual references, or supported short video depending on task. Sound direction can be included in the prompt, allowing a clip to arrive with ambience, effects, and speech. The previous-interaction workflow supports conversational visual edits.
The current model card fixes output at 720p, 24 fps, and 3–10 seconds. Google explicitly says audio references and voice editing are unsupported. A creator cannot supply a voice sample as the audio target or selectively correct the generated speaker while preserving everything else through a documented voice-edit feature.
Omni therefore fits concept work where semantic sound matters more than exact casting. Treat dialogue as provisional until words, pronunciation, voice consent, and timing pass review. The model is in public preview and lacks provisioned throughput, so availability and behavior may change.
Grok Imagine 1.5: a GA image-to-video contract
xAI says Grok Imagine Video 1.5 generates sound effects, ambience, and dialogue in the same pass and improves speech clarity and synchronization over its predecessor. The claim is vendor-reported; each production team should test its languages, shot types, and speech patterns.
The API begins with an image and a motion prompt, then selects duration and 480p, 720p, or 1080p resolution. General availability provides a clearer production status than a preview, while capacity, output variability, and policy still require normal error handling.
The starting frame can establish the speaker and composition, but it does not establish voice rights or vocal identity. Prompt a short line, avoid competing sound instructions, and review lip closure and syllable timing at normal speed and frame-by-frame.
Meta Muse Video: promising, but not a product comparison yet
Meta says Muse Video shares a pretraining base with Muse Image and provides native audio alongside competitive visual fidelity, prompt adherence, and temporal consistency. Its launch article reports an Arena position as of July 5, 2026, but also calls the model an early preview.
Meta explicitly names audio-video synchronization and physically accurate fast motion as areas with current performance gaps. That disclosure is more useful than treating every demo as solved. It suggests evaluations should include quick impacts, rapid speech, occlusion, and scenes where several sounds compete.
No public developer API, price, output resolution, duration, or exact launch date is announced in the source reviewed here. Muse should appear on a research watchlist, not in a committed delivery architecture. Meta also says its Content Seal watermark is planned for video; the existing image implementation should not be described as already active on Muse Video.
Five tests for audiovisual quality
First, test event synchronization. Mark visible contact frames—foot strike, clap, impact, switch, door—and measure whether the sound lands within an acceptable window. Subjective “feels synced” review can miss consistent offset.
Second, test semantic alignment. Rain should not sound like fire, and room ambience should match apparent space. Listen once without picture and watch once muted to isolate each channel.
Third, test dialogue. Transcribe the output, compare it with the approved script, and score intelligibility, pronunciation, language, pacing, and mouth movement. Never assume a plausible sentence is the requested sentence.
Fourth, test mix quality. Measure peaks, integrated loudness, noise, clipping, stereo image, and whether dialogue remains audible over effects. Delivery platforms may impose their own loudness targets.
Fifth, test continuity. If several clips form one scene, voices, room tone, perspective, and musical key should not reset at each generation. Native audio is shot-local unless the product explicitly guarantees sequence continuity.
Consent, rights, and provenance
A generated voice can resemble a real person even without a supplied reference. Review likeness and impersonation risks, especially for ads, political content, healthcare, and financial communication. Obtain consent for any intentional voice identity and prohibit prompts that evade it.
Document the script, audio prompt, model version, generation ID, selected output, and reviewer. Record whether music and effects were generated, replaced, or licensed. If a downstream editor changes the track, update provenance rather than continuing to call it the model’s native audio.
Pricing should also be evaluated audiovisually. Omni output is currently about $0.10 per generated second at 720p before inputs and retries. Grok lists $0.08, $0.14, and $0.25 per output second for 480p, 720p, and 1080p, plus image input. A failed soundtrack can turn a visually accepted generation into another paid attempt.
Preserve, replace, pitch-adjust, or re-sync with Medux
Medux is a separate post-processing service, not the source of Omni, Grok, or Muse native audio. Preserve the model track when dialogue, effects, ambience, rights, and loudness all pass. Otherwise, choose one deliberate correction instead of repeatedly regenerating an approved picture.
The Claude MCP audio-track workflow can attach an approved narration or music file. The Claude MCP pitch-adjustment guide can make a controlled pitch change when appropriate; it must not be used to imitate someone without consent. The Codex MCP lip-sync tutorial addresses visible speech timing after an approved voice track is chosen.
Each operation needs the source hash, approved audio identity, parameters, Medux task ID, status, and output validation. Keep an untouched export. Check duration, stream count, sample rate, synchronization, and playback after processing, and require a person to approve any recognizable voice.
Compare the processed track with the original at matched timecodes and archive both. A successful status only proves that a task completed; it does not prove that words, mouth movement, musical timing, or rights are correct.
This creates a clean decision tree: keep good native sound; replace unusable sound; adjust pitch only for an authorized creative goal; apply lip sync only after the final dialogue is locked. Native audio accelerates ideation, while explicit finishing keeps the published track accountable.