xAI has updated its video model Imagine Video 1.5 with reference images, reference audio, and the ability to generate a clip from a written prompt alone. Output now runs at native 1080p. The goal is to keep the same person or the same product intact from one shot to the next. What used to be a tool for animating a single still image is turning into a tool you brief, shot by shot.

A clip from text alone

Until now, Imagine Video worked mainly from a starting image that the model then set in motion. This update adds a path where a written prompt is enough. Under the hood it chains image generation into image-to-video, but from the user's side, the step of preparing source artwork simply disappears.

Native 1080p arrives alongside it. Both text-to-video and image-to-video are covered, and this part is already generally available. Because 1080p is handled at generation time rather than through upscaling, fine detail should hold up better.

Each reference image pins down one element

Multi-Reference is the centrepiece of the update. Every image passed in as a reference is responsible for a different element. One image holds the face, another preserves the product, another defines the location.

That division of labour is what makes selective edits possible. Keep the character and swap the background, keep the background and swap the character, or hold both and change only the action. A single generation accepts up to 7 references.

Specifying what should not change has been the awkward part of AI video. Teams have been wording and rewording prompts to summon the same character. Now the image itself can serve as the anchor.

Voice references keep face and voice together

The second addition is voice reference. Supply a character image together with a voice sample, and xAI says the model holds both the same face and the same voice across scenes.

For short films or ads built from several shots of one character, rebuilding that character for every cut was the heaviest part of the job. Locking face and voice as a pair reduces the risk that a sequence of separately generated cuts ends up looking like different people. It is a sign that video generation is being judged less on a single beautiful shot and more on continuity across a whole piece.

Staged rollout, with the API ahead

Availability differs by feature.

Feature Availability
Image and voice references Started in the United States for SuperGrok Heavy and SuperGrok Plus on the web and iOS, expanding to all tiers over the following days
Text-to-video Generally available on web, iOS, and Android
Native 1080p Generally available on web, iOS, and Android

Developers can reach the model through the API as grok-imagine-video-1.5. Image references, text-to-video, and native 1080p are all live there, while voice references require a separate request. The contrast is notable: the apps are getting a staged rollout, but the API opens most of the feature set at once.

Imagine Video 1.5 itself launched only the previous month, pitched then on motion, physics, and audio quality. This release layers control on top of that rather than replacing the model, widening what a user is allowed to specify.

Summary

xAI has added up to 7 reference images, voice references, prompt-only generation, and native 1080p output to Imagine Video 1.5. Reference images are designed to pin down individual elements such as a face, a product, or a location, and combining them with a voice reference keeps the same character consistent across multiple shots. The reference features are rolling out from the higher-tier US plans, while text-to-video and 1080p are already generally available. The same model is accessible through the API.