Luma Ray 3.2 and Grok Imagine Video 1.5 arrived one week apart in June 2026, but they are not interchangeable releases. Ray 3.2 is a broad video-generation and transformation system designed around direction and post-production control. Grok Imagine Video 1.5 is a generally available image-to-video model designed to animate one source image with motion and native sound.

That difference matters more than a single leaderboard score. A studio adapting existing footage needs different controls from a retailer turning product photos into social clips. The fairest comparison begins with input, output, and workflow—not with a generic claim that one “wins.”

Capability snapshot

Decision Luma Ray 3.2 Grok Imagine Video 1.5
Release state Public Ray 3.2 API GA in xAI Imagine API
Primary inputs Text, anchor images, or source video One starting image plus text
Generation modes Text-to-video, image-to-video, video editing, extension, reframe Image-to-video
Direction Start/end and indexed keyframes; source-preservation controls Natural-language motion from one image
Audio Not listed as native output Native effects, ambience, and dialogue
Resolution Draft through 1080p, depending on operation 480p, 720p, or 1080p
Duration Current API generation is 5 or 10 seconds; video edits accept source clips up to 18 seconds Current generation docs allow 1–15 seconds
Pro formats HDR MP4 and optional EXR export Standard delivered video
Billing Per request by type, duration, resolution, dynamic range Per output second by resolution, plus image input

The table reflects official documentation available on July 17, 2026. Both providers can change limits and pricing, so applications should not turn these values into permanent UI promises.

Ray 3.2 is a directing and transformation surface

Luma’s strongest differentiator is control over an existing visual plan. The current Agents API supports start and end frames, indexed keyframes, source-video edits, reframing, and extension. It also exposes controls for how strongly motion, structure, pose, and faces carry from source footage.

For multi-shot work, keyframes can anchor a product angle, character pose, camera destination, or important narrative beat. Current Luma documentation lists up to 64 indexed anchors, while the launch announcement and API landing page described up to 16. A production implementation should follow the current schema and test account capabilities rather than relying on either number indefinitely.

Ray 3.2 also targets conventional finishing. Luma lists native HDR generation and 16-bit EXR export, making it possible to carry more image information into grading and compositing. That does not eliminate review or VFX work, but it is a meaningful distinction for teams whose endpoint is Resolve, Nuke, or another managed post pipeline.

Luma’s launch materials marketed video-to-video clips up to 20 seconds. The current Agents API reference is more specific: generated clips use five- or ten-second tiers, while a video-edit source must be 18 seconds or shorter and the output matches its duration. Current schema limits should govern implementation even when a launch headline describes a rounder maximum.

Grok 1.5 begins with a single still and adds sound

The GA model grok-imagine-video-1.5 takes a starting image, a prompt describing action and camera behavior, and settings such as duration and resolution. xAI says the model improves motion, physics, generated audio, and speech synchronization over its previous version.

The sound is generated in the same pass as the video. That is convenient for a concept, meme, product reveal, or short narrative moment because the first preview already has ambience and effects. It is not a guarantee that the audio is legally usable, perfectly synchronized, or appropriate for a brand. Dialogue still needs transcription and review; music and voice-like output need rights and identity controls.

The 1.5 model is narrower than the broader grok-imagine-video family surface. Its official model page lists image input and video output and says it does not support text-to-video. xAI’s reference-to-video documentation also says 1.5 does not support that separate multi-reference mode. Do not assume that every capability documented for the base Imagine video model also exists in 1.5.

What can actually be said about speed

xAI publishes a concrete launch figure for Grok Imagine Video 1.5 Fast: a six-second, 720p clip in about 25 seconds, compared with more than 40 seconds for its previous model. The Fast variant was announced for Grok’s web and mobile Imagine experiences. The GA API model was announced in the same post, but xAI did not state that every API request has the same latency.

Luma’s Ray 3.2 materials describe Speed and Quality inference choices and shared versus provisioned capacity, but the launch page does not provide a directly comparable seconds-to-video figure. It would therefore be misleading to put “25 seconds” beside an invented Ray number.

Benchmark real workloads. Record queue time, generation time, download time, failure rate, and accepted-output rate. A fast rejected clip is slower in business terms than a longer render accepted on its first attempt.

Price comparison needs context

xAI’s current model page lists one image input at $0.01, then output at $0.08 per second for 480p, $0.14 for 720p, and $0.25 for 1080p. A ten-second 720p output is therefore $1.40 plus the image input at listed rates.

Luma’s API page lists SDR text- or image-to-video requests by fixed duration. At the time of writing, ten seconds is approximately $0.45 at 540p, $0.90 at 720p, and $3.60 at 1080p. Video-to-video, HDR, and EXR have different rates.

Those numbers are not a pure value ranking. Grok includes generated audio, while Ray offers keyframe and post-production controls. Compare cost per approved deliverable, including retries, audio replacement, grading, storage, and reviewer time.

Add opportunity cost to the comparison. If a Ray shot requires an editor to prepare several anchors, the generation price understates labor. If a Grok clip needs repeated attempts because one starting image cannot constrain an intermediate beat, its simple request shape understates sampling. Time source preparation, review, and repair separately; the cheapest API row may not produce the cheapest campaign.

Choose Ray 3.2 when control dominates

Ray is the stronger starting point when the workflow needs:

It is also a better conceptual match when a human director wants to settle intermediate beats instead of asking a model to invent the whole path from one image.

Choose Grok 1.5 when a still must become a fast, voiced moment

Grok 1.5 fits product photos, character illustrations, social posts, and storyboard panels that need motion plus sound in one request. Its GA status is useful for production planning, and per-second billing makes duration costs easy to estimate.

It is less suitable when several references must be combined, an existing video must be transformed, or exact intermediate poses are required. In those cases, choose another model rather than forcing a one-image workflow beyond its documented surface.

Run the same acceptance test on both

Use an evaluation set containing portraits, products, typography, fast motion, camera movement, and different aspect ratios. Score:

Run at least one blind review in which evaluators do not know the provider. Then repeat with technical metadata visible so operators can judge controllability and debugging. The first test reduces brand preference; the second captures workflow realities that pure visual scoring misses.

For Ray, add anchor-to-anchor continuity and source-motion preservation. For Grok, add audio relevance, voice consistency, and whether the still remains recognizable throughout the clip.

Normalize downstream finishing

Generated audio is an option, not an obligation. If Grok’s sound is approved, keep it. If it contains unsuitable speech, noise, or music, preserve the video and replace the track. Ray outputs may need an audio track from the outset. Unwanted baked-in subtitles require a visual cleanup decision rather than simply muting audio.

An authorized Codex or Claude MCP client can call Medux after generation, but neither model has a native Medux integration. The audio-track tutorial, subtitle-removal tutorial, and video-merge tutorial illustrate independent finishing operations.

Use the same manifest for both providers: source checksum, model ID, request ID, prompt version, selected output, audio decision, cleanup tasks, and final checksum. Once the evaluation and finishing path is shared, the choice becomes clearer. Ray 3.2 wins when directability and post control drive acceptance; Grok Imagine Video 1.5 wins when a single image, native sound, and rapid ideation define the job.