Gemini Omni Flash and Grok Imagine Video 1.5 arrived in June 2026 with a shared promise: generate short video with sound fast enough for real creative iteration. Their product contracts are not the same. Google emphasizes a unified multimodal model that can create and edit. xAI emphasizes a generally available image-to-video model with resolution choices and faster production.

The right comparison is therefore not a universal quality winner. It is which input contract, output format, availability level, and cost structure fit a specific shot. No independent side-by-side test is claimed here; capabilities and prices reflect official documentation available on July 17, 2026.

The shortest useful comparison

Capability Gemini Omni Flash Grok Imagine Video 1.5
Product status Public preview Generally available in API
Current API focus Text, image, references, and short-video editing Image-to-video
Output resolution 720p 480p, 720p, or 1080p
Duration 3–10 seconds Duration is selected with the request under current API limits
Audio Generated with video Sound effects, ambience, and dialogue
Iteration Previous-interaction editing New image-and-prompt request workflow
Published 720p output price About $0.10 per second $0.14 per second, plus image input

Specifications are only a starting point. Output quality varies with source image, prompt, motion, subject, and sampling. Create a fixed evaluation set before choosing a default.

Gemini Omni Flash favors mixed evidence

Omni’s differentiator is the variety of supported starting points. It can generate from text, animate an image, use tagged image references, or edit supported short video. A creator can continue from a previous interaction and describe a revision conversationally.

That makes it useful when the task is not cleanly “animate this still.” A product team may provide a reference object, a separate environment image, a written camera direction, and then request a change to the generated shot. Keeping those modes under one model reduces format translation between tools.

The current output is constrained to 720p, 24 fps, and 3–10 seconds. Google’s public-preview documentation also lists important exclusions: no audio references, extension, interpolation, voice editing, or reliable use of video references as currently described. Multiple-video reasoning is unavailable. It is a flexible shot tool, not an unrestricted video editor.

Grok Imagine Video 1.5 favors a clear image-to-video contract

xAI launched Grok Imagine Video 1.5 for API use on June 16, 2026. The API model page describes a starting image plus prompt, resolution, and duration. Its output choices cover 480p, 720p, and 1080p, making the contract straightforward for applications that already create or select a hero frame.

xAI says the model improves prompt following, visual quality, motion, and native audio. Its launch announcement reports that a six-second 720p video can be generated in roughly 25 seconds in the Fast experience, compared with more than 40 seconds previously. This is a vendor-reported example, not a universal latency guarantee for every API request or queue condition.

The narrower input mode can be an advantage. A team can approve the exact starting composition before spending on motion. It can also be a limitation when a task needs text-only ideation, several independent references, or conversational modification of an existing video.

Control starts before the prompt

For Grok 1.5, the initial image carries much of the composition, identity, palette, and framing. Prompts can focus on motion, camera, timing, and audio. Poorly resolved anatomy or text in the first frame may persist or move unpredictably, so inspect the image at the target aspect ratio before animation.

For Omni, control can be distributed across a prompt and references. Tag each reference consistently, state what must remain fixed, and avoid assigning incompatible roles to one image. When editing a clip, describe the desired delta rather than restating the entire scene.

Neither approach offers frame-exact control merely because it follows a prompt. Test camera moves, occlusion, hands, faces, text, logos, and object permanence. If an exact frame is legally or commercially required, keep a deterministic post-production route available.

Audio generation is useful but not final by default

Both models can return audio with video. Grok’s launch explicitly describes sound effects, ambience, and dialogue. Omni attempts sound generation from the prompt. This can compress concept development: reviewers hear the intended energy without a separate temporary soundtrack.

Generated audio adds failure modes. Speech can be wrong, accents inconsistent, music inappropriate, or a sound effect out of sync. Neither product’s headline replaces consent for a voice, music rights, loudness normalization, or delivery-format checks. Omni currently does not accept an audio reference or edit an existing voice.

Write audio direction as a separate prompt block so it can be reviewed. Preserve a clean visual master when possible. If the picture is strong but the sound is not, replacing audio is usually more economical than regenerating the entire shot repeatedly.

Availability and operational risk

Gemini Omni Flash is a public-preview model. Its model ID, behavior, quota, supported parameters, regional coverage, and limitations may change. Google does not offer provisioned throughput for it. Applications should pin the preview ID and monitor the API changelog.

Grok Imagine Video 1.5 is generally available in the xAI API, which offers a stronger product-status signal. GA does not mean unlimited capacity or fixed behavior forever. Developers still need async processing, timeouts, error classification, rate handling, and a regression set.

Do not build a single synchronous user request that waits indefinitely for a render. Submit work, preserve the request and returned identifier, poll with backoff, and expose a recoverable status. Cap automatic retries because a failed generation may still be billable or may have created an output before the client timed out.

Current list-price comparison

Google prices Omni input at $1.50 per million tokens and generated video at $17.50 per million output tokens. At the documented 5,792 tokens per video second, output is approximately $0.10 per second for 720p. There is no free tier for the model.

xAI’s model page lists $0.01 for the image input and output at $0.08 per second for 480p, $0.14 for 720p, and $0.25 for 1080p. A six-second output would therefore be about $0.49 at 480p, $0.85 at 720p, or $1.51 at 1080p including one image input, before retries.

Those calculations use current list rates and simple arithmetic. They do not include storage, data transfer, preprocessing, rejected takes, or human work. Compare cost per approved second. A nominally cheaper model can cost more if it needs more attempts.

Choosing by production scenario

Choose Omni when text-only generation, multiple visual references, or conversational editing is central and 720p preview output is sufficient. It is especially attractive for concept iteration that moves between generation and transformation.

Choose Grok when a team already has an approved starting frame, wants a stable image-to-video API contract, values GA status, or needs a 1080p output option. The 480p tier can support inexpensive motion tests before a higher-resolution render, provided the workflow tolerates separate attempts.

For either model, run the same test board: human close-up, hands manipulating an object, product turn, camera move, scene with text, dialogue, and complex ambience. Blind-review prompt adherence, temporal stability, audio usefulness, latency, and accepted-take cost.

A common Medux finishing path

Generator choice does not have to determine the entire pipeline. Export the selected Omni or Grok clip and treat it as a versioned source. Medux is a separate media-processing service; neither official model announcement claims a native Medux connection.

When generated sound fails review, follow the Claude MCP audio-track workflow to add an approved track through an independently configured job. When a sequence contains shots from either generator, the Claude MCP video-merge tutorial provides a defined assembly step.

Record generator, model ID, source-image hash, prompt, resolution, duration, generation job, and selected output. Then record each Medux input hash, parameters, task ID, status, and validation result. Check stream count, duration, aspect ratio, audio presence, and playback before distribution.

Normalize delivery expectations before merging. If the selected clips differ in dimensions, frame rate, orientation, or audio layout, choose one approved target specification and make any conversion an explicit recorded step. Never let an agent guess which version is the campaign master from a similar filename.

This shared finishing layer keeps the creative choice reversible. Omni can serve mixed-input exploration and Grok can animate approved frames, while audio replacement and assembly follow the same auditable process. The best model is the one that produces an acceptable shot efficiently; the best workflow remains the one that can explain exactly how the final file was made.