Gemini Omni Flash and Luma Ray 3.2 solve different versions of the AI filmmaking problem. Google asks how one multimodal model can interpret a creative conversation across text, images, and short video. Luma asks how a filmmaker can direct motion, camera, performance, and visual continuity through many explicit moments.
Both were announced in June 2026, but a feature checklist can obscure their difference. Omni is a flexible generator-editor with native audio ambitions. Ray 3.2 is a visually directed production model with keyframes, video-to-video transformation, longer clips, HDR, and high-end export options. The practical choice depends on whether a project starts with mixed evidence or with a tightly planned visual path.
At-a-glance differences
| Capability | Gemini Omni Flash | Luma Ray 3.2 |
|---|---|---|
| Core approach | Multimodal generation and conversational editing | Directed generation and transformation with keyframes |
| Inputs emphasized | Text, images, tagged references, supported short video | Prompts, source video, keyframes, performance references |
| Published output | 3–10 seconds, 720p, 24 fps | Up to 20 seconds for documented video-to-video workflows, up to 1080p |
| Audio | Attempts audio with the video | Official launch emphasis is visual production control |
| Production formats | Standard generated video with SynthID | HDR and 16-bit EXR options highlighted |
| Current status | Public preview | Available through Luma products and API |
This table compares documented capabilities, not independent output quality. Motion realism and prompt adherence must be tested on a common storyboard.
Omni treats media as conversational context
Gemini Omni Flash supports text-to-video, image-to-video, reference-guided video, and edit tasks under one preview model. A creative can begin with a written idea, anchor it with several visual references, generate a shot, and ask for a revision using the previous interaction.
That workflow suits briefs in which the inputs play different roles. One image may specify a product, another a location, while text defines action, camera, and sound. Tags help the model associate each reference with the intended element. A short source clip can also be transformed within the current editing limits.
Omni’s broad input vocabulary is balanced by narrow output bounds. The model card specifies 720p, 24 fps, and 3–10 seconds. Current documentation excludes extension, interpolation, audio references, voice editing, and multiple-video reasoning. It also warns that video-reference behavior is not yet processed correctly in a documented reference path. Public-preview behavior may change.
Ray 3.2 treats time as a directed structure
Luma presents Ray 3.2 as a model for professional-grade visual direction. Keyframes let a creator specify important moments along a shot instead of relying on one initial image and a prose description. The model fills motion between those anchors while preserving the intended progression.
Luma’s API page states support for up to 16 keyframes in a single clip. Some Luma learning material describes up to 64 keyframes in particular Modify or interface workflows. Those numbers should not be merged into one universal limit. Confirm whether a project uses the API, Dream Machine, Modify, or another surface, and follow the current documentation for that path.
Ray 3.2 also emphasizes performance transfer, including multi-person scenarios, and video-to-video transformation. Luma’s announcement discusses tracking performance across up to eight faces. Such figures are vendor specifications, not a guarantee of identity fidelity under every occlusion, angle, or edit.
Duration and scene construction
Omni generates self-contained short shots. Ten seconds is enough for many ads, transitions, and social moments, but a long sequence must be designed as multiple outputs. Because continuation and interpolation are currently unsupported, continuity between shots needs careful storyboarding and selection.
Ray 3.2’s documented video-to-video path can work with source footage up to 20 seconds, with generated duration tied to the source in relevant workflows. This gives a creator more temporal material and a stronger motion scaffold. It does not mean any text-to-video request automatically returns a 20-second 1080p clip; mode-specific limits still apply.
For both models, avoid asking one shot to do the work of a scene. Break the action around a clear visual beat. Keep camera direction, subject motion, background motion, and transition intent separate. A shorter accepted shot is more valuable than a long result with an unusable middle.
Image quality, HDR, and post-production
Ray 3.2’s visual-production positioning includes 1080p output, HDR, and 16-bit EXR. EXR can preserve a richer image representation for color and compositing workflows than a standard delivery video. This matters when the AI result must enter a conventional post-production pipeline rather than publish directly.
Those formats create operational requirements. Storage grows, playback support differs, and color management must remain consistent. Verify color space, transfer characteristics, alpha behavior where applicable, and how the final delivery encode is produced. An HDR option does not make every display or platform preserve the intended look.
Omni’s 720p output favors speed and iteration. It can be sufficient for concept approval, mobile placements, or workflows that do not require a high-resolution master. Upscaling cannot restore every missing texture or edge, so test the actual delivery size before committing.
Audio changes the comparison
Omni explicitly attempts ambience, effects, and dialogue along with the video. That creates a more complete preview and may help with comedic or rhythmic timing. Current limits prevent supplying an audio reference or selectively editing a generated voice, so a promising visual can still require downstream sound work.
Luma’s official Ray 3.2 materials focus on visual direction, performance, transformation, and professional image output. Do not infer an equivalent audio contract from the word “video.” Plan a separate soundtrack unless the exact Luma product surface and job result prove otherwise.
In either case, review speech, synchronization, rights, loudness, channel layout, and silence. Preserve a picture master so a sound problem does not force visual regeneration.
Pricing and capacity are not directly comparable
Google publishes Omni’s token-based rate. At the documented conversion, 720p output is approximately $0.10 per second, plus input tokens and repeated generations. The model has no free tier and no provisioned-throughput option during the current preview.
Luma’s official API page presents Build and Scale access, but the reviewed Ray 3.2 launch materials do not provide one universal numeric price that can be compared safely with Google’s rate. Current account, plan, resolution, mode, and commercial terms should be checked directly. Do not copy a historic Ray price from a third-party post into a production budget.
Compare cost per approved shot, including source preparation, rejected takes, storage, and finishing. Ray’s higher-end output may avoid another generation stage; Omni’s multimodal iteration may reduce setup time. The cheaper invoice line is not necessarily the cheaper creative process.
Which model fits which job?
Choose Omni when the project begins with a mixture of text and references, when conversational changes are important, or when generated sound accelerates concept review. It is a strong candidate for short, quickly iterated clips where 720p and preview status are acceptable.
Choose Ray 3.2 when the director needs explicit temporal anchors, has source performance to transform, expects a longer video-to-video shot, or needs 1080p, HDR, or EXR-oriented output. Keyframes are particularly valuable when several moments must land in an intentional order.
Test both with the same brief. Score subject continuity, transition accuracy, camera adherence, face and hand stability, text integrity, accepted duration, generation latency, and total attempts. Label vendor-reported features separately from observations made by your team.
A shared Medux finishing workflow
After selection, either model produces an external media asset. There is no claim that Gemini or Luma is natively integrated with Medux. A separately configured Medux MCP workflow can handle deterministic assembly around those approved exports.
Use the Claude MCP video-merge tutorial to combine individually approved shots in a declared order. If Omni’s generated sound is unsuitable, or Ray output needs narration or music, the Claude MCP audio-track guide describes adding a selected track as its own asynchronous job.
Keep provenance through every handoff: model and version, source hashes, prompts or keyframe plan, generation identifier, selected output, Medux input hash, task ID, parameters, and validation. Confirm duration, dimensions, frame rate, color behavior, stream count, and audio after each transformation.
Ray’s HDR or EXR-oriented source deserves an extra gate before a standard delivery merge. Decide where color conversion occurs, preserve the high-quality master, and inspect the encoded result on the intended playback device. For Omni, verify that any generated soundtrack has been deliberately retained or replaced. These decisions should appear in the manifest rather than being inferred from whichever streams happen to be present.
The distinction between the models then remains productive. Omni supplies flexible multimodal exploration; Ray supplies directed visual construction. Medux does not erase those differences—it gives either route a common, auditable way to assemble shots and finish sound.