389| Image-to-video | `prompt` + `image` | `prompt: { image, text }` | Generates video with the provided image as the starting frame. |389| Image-to-video | `prompt` + `image` | `prompt: { image, text }` | Generates video with the provided image as the starting frame. |
390| Reference-to-video | `prompt` + `reference_images` or `reference_audios` | `prompt: "..."` + `providerOptions.xai.{ mode: "reference-to-video", referenceImageUrls }` | Generates video guided by reference images and/or a preset voice on `grok-imagine-video-1.5`. |390| Reference-to-video | `prompt` + `reference_images` or `reference_audios` | `prompt: "..."` + `providerOptions.xai.{ mode: "reference-to-video", referenceImageUrls }` | Generates video guided by reference images and/or a preset voice on `grok-imagine-video-1.5`. |
391| First & Last frame | `last_frame`, optionally with `image` and/or `prompt` | REST body `last_frame` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins the exact last frame. Add `image` to also pin the first frame and interpolate between the two. `prompt` is optional whenever a frame is pinned. Can be combined with `reference_images` / `reference_audios`. |391| First & Last frame | `last_frame`, optionally with `image` and/or `prompt` | REST body `last_frame` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins the exact last frame. Add `image` to also pin the first frame and interpolate between the two. `prompt` is optional whenever a frame is pinned. Can be combined with `reference_images` / `reference_audios`. |
392| Keyframes | `keyframes` (up to 4 `{image, timestamp_s}` entries), optionally with `image`, `last_frame`, and/or `prompt` | REST body `keyframes` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins images at exact moments strictly inside the clip, on a 1/3-second grid. Combine with `image` / `last_frame` for the endpoints and with `reference_images` / `reference_audios` for guidance. |
392| Edit-video | `/v1/videos/edits` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "edit-video", videoUrl }` | Modifies an existing video based on the prompt. |393| Edit-video | `/v1/videos/edits` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "edit-video", videoUrl }` | Modifies an existing video based on the prompt. |
393| Extend-video | `/v1/videos/extensions` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "extend-video", videoUrl }` | Extends an existing video from its last frame. |394| Extend-video | `/v1/videos/extensions` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "extend-video", videoUrl }` | Extends an existing video from its last frame. |
394 395
395396On `grok-imagine-video-1.5`, `image` combined with `reference_images`, `reference_audios`, or `last_frame` is reference-to-video with a pinned first frame. The clip starts on that image rather than treating it as a style reference. `last_frame` on its own (no `image`, no references) is valid: the model generates the opening and lands on the pinned frame. `prompt` is optional for any request that includes `image`, `reference_images`, or `last_frame`; it is required only for text-to-video. Classic `grok-imagine-video` still rejects `last_frame` and rejects combining `image` with reference inputs.On `grok-imagine-video-1.5`, `image` combined with `reference_images`, `reference_audios`, `last_frame`, or `keyframes` is reference-to-video with a pinned first frame. The clip starts on that image rather than treating it as a style reference. `last_frame` or `keyframes` on its own (no `image`, no references) is valid: the model generates the rest of the clip around the pinned frames. `prompt` is optional for any request that includes `image`, `reference_images`, `last_frame`, or `keyframes`; it is required only for text-to-video. Classic `grok-imagine-video` still rejects `last_frame` and `keyframes` and rejects combining `image` with reference inputs.
396 397
397Do not mix AI SDK `mode` values. Each request supports exactly one of `"edit-video"`, `"extend-video"`, or `"reference-to-video"`. When you omit `mode`, the AI SDK uses standard generation.398Do not mix AI SDK `mode` values. Each request supports exactly one of `"edit-video"`, `"extend-video"`, or `"reference-to-video"`. When you omit `mode`, the AI SDK uses standard generation.
398 399
399400See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame) for `last_frame` examples.See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame) for `last_frame` examples and [Keyframes](/developers/model-capabilities/video/reference-to-video#keyframes) for mid-video pins.
400 401
401## Customize Polling Behavior402## Customize Polling Behavior
402 403