Grok Imagine Video 1.5 is xAI’s generally available model for turning one still image into a short video with generated sound. It entered API preview on June 3, 2026, under grok-imagine-video-1.5-preview, then moved to GA on June 16 as grok-imagine-video-1.5.
The release is notable because motion and audio arrive together. A creator can provide a product photograph, illustration, or composed scene; describe camera movement, action, atmosphere, and sound; and receive a video with effects, ambience, or dialogue. That shortens the path to a convincing preview, but it does not remove production review.
The model is image-to-video, not every-to-video
The documented 1.5 model accepts an image and returns video. It does not support text-to-video. It also does not support xAI’s separate reference-to-video mode, which can guide the base Grok Imagine video model with several images.
This distinction is easy to miss because xAI’s broader Imagine API includes text-to-video, image-to-video, reference-to-video, editing, and extension. A capability on the platform is not automatically a capability of every model. For 1.5, begin with one source frame.
The source can be supplied as a public URL, a base64 data URI, or a file_id from xAI’s Files API. The prompt describes how that frame should move rather than restating every visible detail. The output settings include duration and resolution, subject to the current API constraints.
What improved in 1.5
xAI groups the improvements into three areas.
First, motion and physics are intended to hold together for longer. The launch examples emphasize believable weight, momentum, and fewer shape warps. That matters because an image-to-video model must invent what was not visible in the still while preserving what was.
Second, audio is generated in the same pass. xAI describes sound effects, ambience, and dialogue that align with action, plus clearer speech and better synchronization than the prior model. A door can close with an impact, a street can gain environmental sound, and a speaking subject can receive a generated line.
Third, xAI introduced a faster product variant. Grok Imagine Video 1.5 Fast reportedly makes a six-second, 720p video in about 25 seconds, versus more than 40 seconds for the previous model. The announcement places Fast in Grok’s web and mobile Imagine products. Do not convert that figure into an API service-level agreement for the standard GA model.
Resolution, duration, and price
The current model page lists 480p, 720p, and 1080p outputs. xAI’s generation documentation describes configurable short clips, and the launch code uses a ten-second 720p example. Applications should validate duration at request time because platform-level docs and model-specific constraints can evolve.
Pricing is per generated second, plus the source-image input. As of July 17, 2026, xAI lists:
| Resolution | Output price per second |
|---|---|
| 480p | $0.08 |
| 720p | $0.14 |
| 1080p | $0.25 |
The model page lists the image input at $0.01. At those rates, a ten-second 720p clip is $1.40 plus input. That is the request cost, not the cost of an accepted creative. Five attempts, human review, storage, and downstream edits can make the real total several times higher.
Use lower resolution while testing motion and pacing. Promote an approved prompt and source to the final tier. Track cost by source type and acceptance reason so the team can see whether failures come from prompting, the starting composition, audio, or the model itself.
Prompt the change over time
The starting image already defines subject, framing, color, and lighting. A useful prompt focuses on temporal behavior:
- subject action and its intensity;
- camera movement and lens feeling;
- background motion such as smoke, rain, or cloth;
- pacing and the ending state;
- desired sound effects, ambience, and dialogue;
- details that must remain fixed.
For example: “Slow push-in while the bottle rotates ten degrees clockwise; condensation moves naturally; keep the label unchanged; soft studio hum and one restrained glass chime; no speech.” This is more testable than “make an exciting luxury ad.”
Avoid asking one short clip to contain several unrelated scenes. Stage separate source images and generate separate shots when the camera, location, or narrative state changes. That makes continuity and retries easier to control.
Native audio is a draft and a deliverable candidate
Generated sound can be good enough to keep, especially for low-risk ambience or an illustrative effect. It can also fail independently from the video. A strong visual clip may contain clipped speech, an unwanted voice, mistimed impact, distracting music, or semantic audio that contradicts the action.
Review audio on its own:
- transcribe all intelligible speech;
- verify language, names, numbers, and claims;
- check synchronization at important impacts and mouth movements;
- inspect loudness, clipping, noise, and channel layout;
- confirm that voices, music, and references are authorized;
- decide whether to keep, mute, or replace the track.
Do not assume a generated voice has consent because it came from a model. Avoid prompts that imitate a real person without authorization, and require heightened review for ads, politics, health, finance, or public-safety content.
GA does not mean deterministic
General availability signals a supported production model rather than an experimental preview, but generative outputs remain variable. The same prompt and image can produce different motion, geometry, or audio. New backend versions may also change behavior.
The official model page currently lists the old preview name and a dated name as aliases. For stable operations, record the requested model, returned model or fingerprint when exposed, prompt version, input checksum, duration, resolution, request ID, and acceptance result. Maintain a regression set instead of judging a model update from one attractive clip.
Generated URLs can be ephemeral. Download accepted outputs promptly or use xAI’s file-output controls. A job is complete only after the asset is stored, checksummed, and linked to its metadata.
Create a rejection taxonomy before a large run. Useful labels include source drift, geometry, identity, text, motion, physics, camera, audio relevance, dialogue, synchronization, policy, and technical delivery. Reviewers should select the dominant failure and may add secondary labels. Over dozens of jobs, those counts reveal whether a prompt template needs revision, a source composition is unsuitable, or the model should be routed away from a particular task.
Retain rejected outputs only as long as governance and evaluation require. They may contain unwanted identities, text, or audio and should not be casually exposed in a shared creative library. Separate review storage from approved publishing storage, and prevent automated systems from selecting an asset merely because it is the latest file.
Best fits and poor fits
Grok Imagine Video 1.5 is a strong fit for:
- animating approved product hero images;
- turning illustrations into social clips;
- testing camera movement from storyboard panels;
- creating short atmospheric moments with sound;
- producing several motion concepts from one composed frame.
Choose another workflow when the task needs text-to-video with no source image, several reference images, surgical editing of an existing video, exact multi-keyframe choreography, or long-form continuity. A model’s narrow input contract can be an advantage when it matches the job and a source of waste when it does not.
The starting image itself can encode more control through conventional design. A compositor may place the product, background, and lighting into one licensed source frame before animation. That remains a single-image request, but the composition should be disclosed and versioned so reviewers know which details came from source design and which were invented during video generation.
From native output to finished asset
After approval, treat audio, branding, and sequence assembly as separate operations. If native sound is suitable, preserve it and document the decision. If it is not, replace the track without regenerating an otherwise accepted clip. Add a logo only after composition is locked, and append an end card as a separate approved clip.
An authorized Codex or Claude MCP client can call Medux for those downstream tasks; Grok Imagine does not natively integrate with Medux. The audio-track workflow demonstrates signed upload, asynchronous task creation, polling, and result download. The Codex logo tutorial and Claude video-merge tutorial cover independent branding and end-card assembly.
Store the xAI request ID and Medux task IDs in one manifest, but do not merge their retry logic. A failed audio replacement should resume the Medux task or retry that step; it should not silently pay for another Grok generation. Grok Imagine Video 1.5 can provide a remarkably complete first clip, while production quality comes from preserving the parts that worked and changing only the parts that did not.