SpyBara
Go Premium

model-capabilities/video/reference-to-video.md 2026-09-25 23:01 UTC to 2026-09-26 23:59 UTC

This page contains 25 additions and 76 deletions.

2026
Tue 8 21:00 Tue 22 23:58 Sat 26 23:59 Mon 28 23:57 Tue 29 23:59

Model Capabilities

Reference-to-Video

Provide reference images, a preset voice, or both to guide the generated video. Images incorporate specific people, objects, clothing, or other visual elements without locking the first frame (unlike image-to-video). This is useful for virtual try-on, product placement, character-consistent storytelling, and voice identity. On grok-imagine-video-1.5, you can also pick the voice your subject speaks in (see Reference audio), pin the exact first or last frame (see First & Last frame), and pin frames at chosen moments inside the clip (see Keyframes).

Each reference image can be provided as a public HTTPS URL, a base64-encoded data URI, or a file_id from the Files API — and you can mix kinds within a single request. See Imagine → Files API Integration for file_id details and examples.

In the Vercel AI SDK, set providerOptions.xai.mode to "reference-to-video" and pass the images with providerOptions.xai.referenceImageUrls.

[!WARNING]

import os
import xai_sdk

client = xai_sdk.Client(api_key=os.getenv("XAI_API_KEY"))

response = client.video.generate(
    prompt="slow zoom in on the white fashion runway stage. then, the model from <IMAGE_1> walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from <IMAGE_2> and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile.",
    model="grok-imagine-video-1.5",
    reference_image_urls=[
        "<IMAGE_URL_1>",
        "<IMAGE_URL_2>",
        "<IMAGE_URL_3>",
    ],
    duration=10,
    aspect_ratio="16:9",
    resolution="720p",
)

print(response.url)
import os
import time
import requests

headers = {
    "Content-Type": "application/json",
    "Authorization": f"Bearer {os.environ['XAI_API_KEY']}",
}

response = requests.post(
    "https://api.x.ai/v1/videos/generations",
    headers=headers,
    json={
        "model": "grok-imagine-video-1.5",
        "prompt": "slow zoom in on the white fashion runway stage. then, the model from <IMAGE_1> walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from <IMAGE_2> and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile.",
        "reference_images": [
            {"url": "<IMAGE_URL_1>"},
            {"url": "<IMAGE_URL_2>"},
            {"url": "<IMAGE_URL_3>"},
        ],
        "duration": 10,
        "aspect_ratio": "16:9",
        "resolution": "720p",
    },
)

request_id = response.json()["request_id"]

while True:
    result = requests.get(
        f"https://api.x.ai/v1/videos/{request_id}",
        headers={"Authorization": headers["Authorization"]},
    )
    data = result.json()
    if data["status"] == "done":
        print(data["video"]["url"])
        break
    elif data["status"] == "expired":
        print("Request expired")
        break
    time.sleep(5)
import { xai } from "@ai-sdk/xai";
import { experimental_generateVideo as generateVideo } from "ai";

const result = await generateVideo({
    model: xai.video("grok-imagine-video-1.5"),
    prompt: "slow zoom in on the white fashion runway stage. then, the model from <IMAGE_1> walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from <IMAGE_2> and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile.",
    duration: 10,
    aspectRatio: "16:9",
    providerOptions: {
        xai: {
            mode: "reference-to-video",
            referenceImageUrls: [
                "<IMAGE_URL_1>",
                "<IMAGE_URL_2>",
                "<IMAGE_URL_3>",
            ],
            resolution: "720p",
            pollTimeoutMs: 600000,
        },
    },
});

const videoUrl = result.providerMetadata?.xai?.videoUrl;
console.log(videoUrl);
# Start the reference-to-video request
REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-imagine-video-1.5",
    "prompt": "slow zoom in on the white fashion runway stage. then, the model from <IMAGE_1> walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from <IMAGE_2> and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile.",
    "reference_images": [
      {"url": "<IMAGE_URL_1>"},
      {"url": "<IMAGE_URL_2>"},
      {"url": "<IMAGE_URL_3>"}
    ],
    "duration": 10,
    "aspect_ratio": "16:9",
    "resolution": "720p"
  }' | jq -r '.request_id')

# Poll until the video is ready
while true; do
  RESULT=$(curl -s https://api.x.ai/v1/videos/$REQUEST_ID \
    -H "Authorization: Bearer $XAI_API_KEY")
  STATUS=$(echo "$RESULT" | jq -r '.status')
  if [ "$STATUS" = "done" ]; then
    echo "$RESULT" | jq -r '.video.url'
    break
  elif [ "$STATUS" = "failed" ] || [ "$STATUS" = "expired" ]; then
    echo "Request $STATUS"; echo "$RESULT" | jq .
    break
  fi
  sleep 5
done

Reference audio

[!WARNING]

Preset voices are generally available. Voice references with your own audio files are available to trusted partners, on request. .

On grok-imagine-video-1.5, give your subject a voice by passing up to 3 preset voices with reference_audios. Each entry names a voice by voice_id, drawn from the same built-in roster as Text to Speech, so {"voice_id": "eve"} speaks in Eve's voice. Identifiers are case-insensitive; an unknown one returns 400 with the list of available voices. You can hear every voice in the flagship voices announcement.

reference_audios accepts preset voices; voice references with your own audio files are available to trusted partners on request. Use a voice alongside reference images or on its own, and tag voices in the prompt as <AUDIO_0>, <AUDIO_1>, and <AUDIO_2> (with <IMAGE_0>… when you also pass images).

import os
import xai_sdk

client = xai_sdk.Client(api_key=os.getenv("XAI_API_KEY"))

response = client.video.generate(
    prompt="The person from <IMAGE_1> presents the product from <IMAGE_2> on the set from <IMAGE_3>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
    model="grok-imagine-video-1.5",
    reference_image_urls=[
        "<IMAGE_URL_1>",
        "<IMAGE_URL_2>",
        "<IMAGE_URL_3>",
    ],
    reference_audios=[
        {"voice_id": "eve"},
        {"voice_id": "leo"},
    ],
    duration=8,
    aspect_ratio="9:16",
    resolution="720p",
)

print(response.url)
import os
import time
import requests

headers = {
    "Content-Type": "application/json",
    "Authorization": f"Bearer {os.environ['XAI_API_KEY']}",
}

response = requests.post(
    "https://api.x.ai/v1/videos/generations",
    headers=headers,
    json={
        "model": "grok-imagine-video-1.5",
        "prompt": "The person from <IMAGE_1> presents the product from <IMAGE_2> on the set from <IMAGE_3>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
        "reference_images": [
            {"url": "<IMAGE_URL_1>"},
            {"url": "<IMAGE_URL_2>"},
            {"url": "<IMAGE_URL_3>"},
        ],
        "reference_audios": [
            {"voice_id": "eve"},
            {"voice_id": "leo"},
        ],
        "duration": 8,
        "aspect_ratio": "9:16",
        "resolution": "720p",
    },
)

request_id = response.json()["request_id"]

while True:
    result = requests.get(
        f"https://api.x.ai/v1/videos/{request_id}",
        headers={"Authorization": headers["Authorization"]},
    )
    data = result.json()
    if data["status"] == "done":
        print(data["video"]["url"])
        break
    elif data["status"] == "expired":
        print("Request expired")
        break
    time.sleep(5)
REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-imagine-video-1.5",
    "prompt": "The person from <IMAGE_1> presents the product from <IMAGE_2> on the set from <IMAGE_3>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
    "reference_images": [
      {"url": "<IMAGE_URL_1>"},
      {"url": "<IMAGE_URL_2>"},
      {"url": "<IMAGE_URL_3>"}
    ],
    "reference_audios": [
      {"voice_id": "eve"},
      {"voice_id": "leo"}
    ],
    "duration": 8,
    "aspect_ratio": "9:16",
    "resolution": "720p"
  }' | jq -r '.request_id')

while true; do
  RESULT=$(curl -s https://api.x.ai/v1/videos/$REQUEST_ID \
    -H "Authorization: Bearer $XAI_API_KEY")
  STATUS=$(echo "$RESULT" | jq -r '.status')
  if [ "$STATUS" = "done" ]; then
    echo "$RESULT" | jq -r '.video.url'
    break
  elif [ "$STATUS" = "failed" ] || [ "$STATUS" = "expired" ]; then
    echo "Request $STATUS"; echo "$RESULT" | jq .
    break
  fi
  sleep 5
done

First & Last frame

On grok-imagine-video-1.5, last_frame pins the exact last frame of the clip. The video ends arriving on that image rather than re-rendering it as a reference. image combined with reference_images, reference_audios, or last_frame is the matching first-frame pin.

Request shape Result
image + last_frame Pinned first and last frame. The model interpolates between the two.
last_frame only Pinned last frame. The model generates the opening and lands on the pinned image.
last_frame + reference_images / reference_audios Pinned last frame with reference guidance. Add image to pin the first frame as well.

prompt is optional in every first & last frame request. Include one to steer motion and camera work between the frames; omit it to let the frames alone drive the clip.

last_frame uses the same URL, data-URI, and file_id shapes as image-to-video. In the xAI Python SDK, pass last_frame_url (or last_frame_file_id for a Files API upload) alongside image_url (or image_file_id). The Vercel AI SDK does not expose last_frame yet; send it on the REST body.

Classic grok-imagine-video rejects last_frame and rejects combining image with reference inputs.

import os
import xai_sdk

client = xai_sdk.Client(api_key=os.getenv("XAI_API_KEY"))

response = client.video.generate(
    prompt="The camera dollies from the sunlit doorway to the window, settling on the closing frame.",
    model="grok-imagine-video-1.5",
    image_url="<FIRST_FRAME_URL>",
    last_frame_url="<LAST_FRAME_URL>",
    duration=8,
    aspect_ratio="16:9",
    resolution="720p",
)

print(response.url)
import os
import time
import requests

headers = {
    "Content-Type": "application/json",
    "Authorization": f"Bearer {os.environ['XAI_API_KEY']}",
}

response = requests.post(
    "https://api.x.ai/v1/videos/generations",
    headers=headers,
    json={
        "model": "grok-imagine-video-1.5",
        "prompt": "The camera dollies from the sunlit doorway to the window, settling on the closing frame.",
        "image": {"url": "<FIRST_FRAME_URL>"},
        "last_frame": {"url": "<LAST_FRAME_URL>"},
        "duration": 8,
        "aspect_ratio": "16:9",
        "resolution": "720p",
    },
)

request_id = response.json()["request_id"]

while True:
    result = requests.get(
        f"https://api.x.ai/v1/videos/{request_id}",
        headers={"Authorization": headers["Authorization"]},
    )
    data = result.json()
    if data["status"] == "done":
        print(data["video"]["url"])
        break
    elif data["status"] == "expired":
        print("Request expired")
        break
    time.sleep(5)
REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-imagine-video-1.5",
    "prompt": "The camera dollies from the sunlit doorway to the window, settling on the closing frame.",
    "image": {"url": "<FIRST_FRAME_URL>"},
    "last_frame": {"url": "<LAST_FRAME_URL>"},
    "duration": 8,
    "aspect_ratio": "16:9",
    "resolution": "720p"
  }' | jq -r '.request_id')

while true; do
  RESULT=$(curl -s https://api.x.ai/v1/videos/$REQUEST_ID \
    -H "Authorization: Bearer $XAI_API_KEY")
  STATUS=$(echo "$RESULT" | jq -r '.status')
  if [ "$STATUS" = "done" ]; then
    echo "$RESULT" | jq -r '.video.url'
    break
  elif [ "$STATUS" = "failed" ] || [ "$STATUS" = "expired" ]; then
    echo "Request $STATUS"; echo "$RESULT" | jq .
    break
  fi
  sleep 5
done

Keyframes

On grok-imagine-video-1.5, keyframes pins images at chosen moments inside the clip. Each entry pairs an image with a timestamp_s, and the video passes through that exact image at that time. Use it to storyboard a shot: the model generates the motion between your anchors rather than inventing the whole clip from one frame.

"keyframes": [
  {"image": {"url": "<KEYFRAME_URL_1>"}, "timestamp_s": 2.0},
  {"image": {"url": "<KEYFRAME_URL_2>"}, "timestamp_s": 4.0}
]

Keyframes cover the interior of the clip; the endpoints keep their own fields. Pin the opening with image and the closing with last_frame, and combine any of the three with reference_images or reference_audios. prompt is optional whenever a frame is pinned.

Constraint Detail
Count At most 4 keyframes per request.
Timing timestamp_s must fall strictly inside the clip: greater than 0 and less than duration. Use image / last_frame for the endpoints.
Spacing Anchors snap to a 1/3-second grid. Two keyframes that round to the same slot are rejected, so keep them at least 1/3 s apart.
Inputs Each image accepts the same URL, data-URI, and file_id shapes as image-to-video.

In the xAI Python SDK, keyframes is a list of dicts that flatten the REST image object: each entry takes image_url (or image_file_id) and timestamp, the REST timestamp_s in seconds, as in {"image_url": "<KEYFRAME_URL_1>", "timestamp": 2.0}. Pin the endpoints with image_url and last_frame_url. The Vercel AI SDK does not expose keyframes yet; send it on the REST body.

Classic grok-imagine-video rejects keyframes, and keyframes cannot be combined with video editing.

Example: a year in one request

This request uses every kind of pin at once. A single oak is pinned as winter in the first frame, spring and summer as keyframes at 3 and 6 seconds, and autumn as the last frame at 9 seconds; the model animates the seasons in between. Select a pin to see the exact frame the video passes through, or press play to watch the whole year.

Pins work best when they share a scene. The four stills started as one generated winter image, and each season is an image edit of it, so the tree, horizon, and camera angle stay identical and the model only has to animate the change of season.