Luma Ray 3.2 changes the control vocabulary of image-to-video generation. Instead of asking a model to infer an entire shot from one starting image and a prompt, a creator can place up to 16 keyframes inside a clip and describe a more detailed sequence of visual states.
Luma released Ray 3.2 on June 9, 2026 for professional production and made the model’s broader control surface available through an API. The launch also highlights enhanced performance tracking, native HDR generation, 16-bit EXR export, Reframe operations, and clips up to 20 seconds at 1080p.
The headline number does not mean a generated clip becomes conventional keyframe animation. In a timeline editor, an animator can expect a parameter to hit a precisely defined value at a precisely defined frame. In a generative model, visual keyframes constrain a learned synthesis process. They give the model much more direction, but the transitions, timing, geometry, and fine motion still need review.
Why 16 keyframes are a meaningful change
Traditional image-to-video commonly starts from one image. The prompt asks for camera movement, subject action, and atmosphere, while the model invents what happens next. Start-and-end-frame workflows add a destination, but a long interval can still wander between those anchors.
Up to 16 keyframes let a director express intermediate beats. A product can begin in a wide shot, rotate toward camera, reveal one feature, move into a close-up, and end in a clean pack shot. A character can enter, react, turn, and reach a final pose. A camera path can pass through several planned compositions rather than relying on one sentence.
More anchors can reduce ambiguity, especially when a client has already approved a storyboard. They can also create conflicting constraints. If adjacent frames change identity, lighting, lens perspective, or object geometry too dramatically, the model must invent an implausible transition. Keyframes should describe a coherent shot, not sixteen unrelated images.
Use the minimum number that communicates the necessary beats. Reserve dense placement for changes that matter: pose, camera position, product orientation, reveal timing, or environmental transformation. A keyframe for every tiny variation can overconstrain motion and make revisions harder.
Designing a multi-keyframe shot
Begin with a short shot specification: subject, story purpose, duration, aspect ratio, camera behavior, non-negotiable details, and the emotion or product fact the viewer should understand. Then choose the major beats before generating images.
Keep identity and production design stable across reference frames. Use the same character features, wardrobe, product proportions, lighting direction, and environment logic. If the sequence changes time of day or location, make that transition intentional rather than accidental.
Treat spacing as an editorial choice. Closely spaced keyframes imply a fast progression; widely spaced anchors give the model more room to interpolate. Luma’s interface or API parameters should be the authority for exact timing behavior, so record the positions used for every approved render.
Draft at lower cost where possible. Review whether the beat order reads correctly before requesting premium resolution or HDR. A high-dynamic-range render will not repair an incoherent action.
Finally, check the transition regions rather than only the supplied frames. Artifacts often appear between anchors: fingers merge, a logo changes, background geometry bends, an object disappears, or a face subtly becomes another person. The control frames are inputs, not a guarantee that every intervening frame is valid.
Performance tracking and expressive faces
Luma says Ray 3.2 improves performance tracking so skeletal posture and gestures can carry into generated video. It also advertises expressive facial performance tracking for up to eight faces at once.
That capability targets scenes where motion is more important than a generic animation: a performer’s stance, timing, head movement, or group reaction. It can support stylized character transformations, virtual production experiments, and localized or alternate visual treatments based on an authorized performance.
“Up to eight faces” is a technical capacity, not a promise that every face remains perfectly identified through occlusion, profile turns, fast motion, or difficult light. Inspect each participant frame by frame. Look for identity drift, gaze errors, expression mismatch, mouth artifacts, and unintended blending between nearby people.
Use only performances and identities that the project is authorized to transform. Obtain consent for the intended output and distribution, document any synthetic alteration, and avoid implying an endorsement that the depicted person did not make. Stronger tracking increases the need for governance because it can preserve a recognizable performance more convincingly.
HDR and 16-bit EXR for post-production
Ray 3.2 supports native HDR generation and 16-bit EXR export. These features are designed for color grading and compositing workflows that need more range and precision than a compressed preview file.
EXR is useful when a VFX or finishing team wants to integrate a generated element alongside live-action plates, manage highlights, or perform heavier grading. The format does not make a shot automatically production-ready. Teams still need to confirm color space, transfer characteristics, alpha behavior if relevant, frame rate, compression, and how the generated image responds to grading.
Review HDR on a managed display and create appropriate SDR deliverables. A bright preview on an unmanaged screen does not validate highlight detail. Preserve the original high-quality output and derive distribution versions rather than repeatedly transcoding one compressed file.
Luma’s current API page says HDR output costs twice the listed SDR price, while HDR plus EXR costs three times the SDR price. Use those multipliers only as a current planning reference and verify the account’s live quote before a batch job.
Reframe and longer shots
Luma presents Reframe as a way to change aspect ratio, extend a frame, or replace a background while preserving the source lighting. This can help adapt a hero shot into vertical, square, and widescreen placements without regenerating the entire concept.
Reframing is more than adding pixels around a subject. Vertical composition may need a different crop, subject scale, eyeline, and safe area for interface overlays. Background replacement must preserve shadows, reflections, contact points, and color spill. Treat Luma’s preservation language as a product claim to evaluate on the actual footage.
Ray 3.2 also supports clips up to 20 seconds at 1080p, with the API page specifically describing video-to-video at up to 20 seconds and 1080p output across the model. Longer duration gives a scene time to develop, but it multiplies opportunities for drift. A production may still be more reliable as several reviewed shots assembled later.
API access and current pricing
Luma says the Ray 3.2 API exposes the full control surface, including Multi-Keyframe. That makes the model usable inside storyboard tools, render pipelines, campaign systems, and internal review applications rather than only through a consumer interface.
Production integration should treat each render as an asynchronous job. Store request IDs, input checksums, model and parameter versions, timestamps, estimated cost, and returned assets. Use idempotency or a local submission ledger so a network timeout does not create a duplicate paid render. Apply rate-aware queues rather than launching an unlimited batch.
As of July 17, 2026, Luma’s public pay-as-you-go page lists approximate text-to-video or image-to-video prices for five seconds at $0.15 for 540p, $0.30 for 720p, and $1.20 for 1080p. Ten seconds is listed at $0.45, $0.90, and $3.60 respectively. The page says there is no minimum commitment, rate limits apply, and the self-serve tier has no latency service-level agreement.
These prices are approximate, based on billing tokens, and can change. Video-to-video, Reframe, HDR, and EXR have different rates. Estimate cost from the exact operation and count rejected renders, retries, storage, and review—not just the final clip.
Best uses and practical limits
Ray 3.2 is a strong fit for storyboard-driven commercials, product reveals, music-video shots, game cinematics, performance transformation, and sequences that need several planned beats. Its high-end output options suit teams that already have color and compositing expertise.
It is less compelling when the task is a simple deterministic edit, when a long scene needs exact frame-level physics, or when legal and identity constraints make generative variation unacceptable. Keyframes improve direction; they do not turn probabilistic synthesis into conventional rendering.
Build an evaluation set around the project’s hardest material: multiple faces, hands touching products, readable packaging, reflective surfaces, camera occlusion, rapid motion, and transitions between distant keyframes. Compare usable-shot rate and revision time, not a curated demo.
Finishing controlled shots with Medux
Once individual Ray 3.2 shots pass review, the production may need to join them, attach approved sound, or add a supplied brand mark. Those are downstream operations rather than reasons to regenerate the visuals.
Medux offers separate media-processing tools that Codex or Claude can use after independent MCP setup. There is no claimed native Luma-to-Medux integration. The Claude video-merge tutorial covers assembling clips in a specified order, the Claude audio-track tutorial shows how to attach selected audio, and the Codex logo tutorial applies a supplied logo to video.
A practical sequence is:
- generate each shot with recorded keyframes and Ray settings;
- review continuity, identity, products, rights, and technical properties;
- export approved masters to controlled storage;
- merge shots in the locked editorial order;
- add licensed audio and approved branding as separate, auditable steps;
- inspect the final duration, streams, sync, dimensions, and visible mark.
Codex or Claude can coordinate these tasks, but it should retain each service’s task ID, poll rather than blindly resubmit, and keep Luma generation charges separate from Medux processing credits. Human approval should remain between creative generation and final publication.
Ray 3.2’s real advance is not the promise of perfect control. It gives creators a denser language for communicating a shot to a generative model. When those anchors are coherent, outputs are tested between keyframes, and finishing remains explicit, multi-keyframe image-to-video becomes a practical production instrument rather than a one-prompt novelty.