2 2
3# Reference-to-Video3# Reference-to-Video
4 4
5Provide reference images, a preset voice, or both to guide the generated video. Images incorporate specific people, objects, clothing, or other visual elements without locking the first frame (unlike [image-to-video](/developers/model-capabilities/video/image-to-video)). This is useful for virtual try-on, product placement, character-consistent storytelling, and voice identity. On `grok-imagine-video-1.5`, you can also pick the voice your subject speaks in (see [Reference audio](#reference-audio)).5Provide reference images, a preset voice, or both to guide the generated video. Images incorporate specific people, objects, clothing, or other visual elements without locking the first frame (unlike [image-to-video](/developers/model-capabilities/video/image-to-video)). This is useful for virtual try-on, product placement, character-consistent storytelling, and voice identity. On `grok-imagine-video-1.5`, you can also pick the voice your subject speaks in (see [Reference audio](#reference-audio)), and pin the exact first or last frame (see [First & Last frame](#first--last-frame)).
6 6
7Each reference image can be provided as a public HTTPS URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.7Each reference image can be provided as a public HTTPS URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.
8 8
257done257done
258```258```
259 259
260## First & Last frame
261
262On `grok-imagine-video-1.5`, `last_frame` pins the exact last frame of the clip. The video ends arriving on that image rather than re-rendering it as a reference. `image` combined with `reference_images`, `reference_audios`, or `last_frame` is the matching first-frame pin.
263
264| Request shape | Result |
265|---------------|--------|
266| `image` + `last_frame` | Pinned first and last frame. The model interpolates between the two. |
267| `last_frame` only | Pinned last frame. The model generates the opening and lands on the pinned image. |
268| `last_frame` + `reference_images` / `reference_audios` | Pinned last frame with reference guidance. Add `image` to pin the first frame as well. |
269
270`prompt` is optional in every first & last frame request. Include one to steer motion and camera work between the frames; omit it to let the frames alone drive the clip.
271
272`last_frame` uses the same URL, data-URI, and `file_id` shapes as [image-to-video](/developers/model-capabilities/video/image-to-video). The Python SDK and Vercel AI SDK do not yet expose a dedicated `last_frame` parameter; send it on the REST body.
273
274Classic `grok-imagine-video` rejects `last_frame` and rejects combining `image` with reference inputs.
275
276```python customLanguage="pythonRequests"
277import os
278import time
279import requests
280
281headers = {
282 "Content-Type": "application/json",
283 "Authorization": f"Bearer {os.environ['XAI_API_KEY']}",
284}
285
286response = requests.post(
287 "https://api.x.ai/v1/videos/generations",
288 headers=headers,
289 json={
290 "model": "grok-imagine-video-1.5",
291 "prompt": "The camera dollies from the sunlit doorway to the window, settling on the closing frame.",
292 "image": {"url": "<FIRST_FRAME_URL>"},
293 "last_frame": {"url": "<LAST_FRAME_URL>"},
294 "duration": 8,
295 "aspect_ratio": "16:9",
296 "resolution": "720p",
297 },
298)
299
300request_id = response.json()["request_id"]
301
302while True:
303 result = requests.get(
304 f"https://api.x.ai/v1/videos/{request_id}",
305 headers={"Authorization": headers["Authorization"]},
306 )
307 data = result.json()
308 if data["status"] == "done":
309 print(data["video"]["url"])
310 break
311 elif data["status"] == "expired":
312 print("Request expired")
313 break
314 time.sleep(5)
315```
316
317```bash
318REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
319 -H "Content-Type: application/json" \
320 -H "Authorization: Bearer $XAI_API_KEY" \
321 -d '{
322 "model": "grok-imagine-video-1.5",
323 "prompt": "The camera dollies from the sunlit doorway to the window, settling on the closing frame.",
324 "image": {"url": "<FIRST_FRAME_URL>"},
325 "last_frame": {"url": "<LAST_FRAME_URL>"},
326 "duration": 8,
327 "aspect_ratio": "16:9",
328 "resolution": "720p"
329 }' | jq -r '.request_id')
330
331while true; do
332 RESULT=$(curl -s https://api.x.ai/v1/videos/$REQUEST_ID \
333 -H "Authorization: Bearer $XAI_API_KEY")
334 STATUS=$(echo "$RESULT" | jq -r '.status')
335 if [ "$STATUS" = "done" ]; then
336 echo "$RESULT" | jq -r '.video.url'
337 break
338 elif [ "$STATUS" = "failed" ] || [ "$STATUS" = "expired" ]; then
339 echo "Request $STATUS"; echo "$RESULT" | jq .
340 break
341 fi
342 sleep 5
343done
344```
345
260## Related346## Related
261 347
262* [Video Generation](/developers/model-capabilities/video/generation) — Generate videos from text prompts348* [Video Generation](/developers/model-capabilities/video/generation) — Generate videos from text prompts