SpyBara
Go Premium

Documentation 2026-09-21 23:00 UTC to 2026-09-22 23:58 UTC

7 files changed +158 −12. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03
Details

389| Image-to-video | `prompt` + `image` | `prompt: { image, text }` | Generates video with the provided image as the starting frame. |389| Image-to-video | `prompt` + `image` | `prompt: { image, text }` | Generates video with the provided image as the starting frame. |

390| Reference-to-video | `prompt` + `reference_images` or `reference_audios` | `prompt: "..."` + `providerOptions.xai.{ mode: "reference-to-video", referenceImageUrls }` | Generates video guided by reference images and/or a preset voice on `grok-imagine-video-1.5`. |390| Reference-to-video | `prompt` + `reference_images` or `reference_audios` | `prompt: "..."` + `providerOptions.xai.{ mode: "reference-to-video", referenceImageUrls }` | Generates video guided by reference images and/or a preset voice on `grok-imagine-video-1.5`. |

391| First & Last frame | `last_frame`, optionally with `image` and/or `prompt` | REST body `last_frame` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins the exact last frame. Add `image` to also pin the first frame and interpolate between the two. `prompt` is optional whenever a frame is pinned. Can be combined with `reference_images` / `reference_audios`. |391| First & Last frame | `last_frame`, optionally with `image` and/or `prompt` | REST body `last_frame` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins the exact last frame. Add `image` to also pin the first frame and interpolate between the two. `prompt` is optional whenever a frame is pinned. Can be combined with `reference_images` / `reference_audios`. |

392| Keyframes | `keyframes` (up to 4 `{image, timestamp_s}` entries), optionally with `image`, `last_frame`, and/or `prompt` | REST body `keyframes` (no dedicated AI SDK field) | On `grok-imagine-video-1.5`, pins images at exact moments strictly inside the clip, on a 1/3-second grid. Combine with `image` / `last_frame` for the endpoints and with `reference_images` / `reference_audios` for guidance. |

392| Edit-video | `/v1/videos/edits` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "edit-video", videoUrl }` | Modifies an existing video based on the prompt. |393| Edit-video | `/v1/videos/edits` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "edit-video", videoUrl }` | Modifies an existing video based on the prompt. |

393| Extend-video | `/v1/videos/extensions` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "extend-video", videoUrl }` | Extends an existing video from its last frame. |394| Extend-video | `/v1/videos/extensions` + `video` | `prompt: "..."` + `providerOptions.xai.{ mode: "extend-video", videoUrl }` | Extends an existing video from its last frame. |

394 395 

395On `grok-imagine-video-1.5`, `image` combined with `reference_images`, `reference_audios`, or `last_frame` is reference-to-video with a pinned first frame. The clip starts on that image rather than treating it as a style reference. `last_frame` on its own (no `image`, no references) is valid: the model generates the opening and lands on the pinned frame. `prompt` is optional for any request that includes `image`, `reference_images`, or `last_frame`; it is required only for text-to-video. Classic `grok-imagine-video` still rejects `last_frame` and rejects combining `image` with reference inputs.396On `grok-imagine-video-1.5`, `image` combined with `reference_images`, `reference_audios`, `last_frame`, or `keyframes` is reference-to-video with a pinned first frame. The clip starts on that image rather than treating it as a style reference. `last_frame` or `keyframes` on its own (no `image`, no references) is valid: the model generates the rest of the clip around the pinned frames. `prompt` is optional for any request that includes `image`, `reference_images`, `last_frame`, or `keyframes`; it is required only for text-to-video. Classic `grok-imagine-video` still rejects `last_frame` and `keyframes` and rejects combining `image` with reference inputs.

396 397 

397Do not mix AI SDK `mode` values. Each request supports exactly one of `"edit-video"`, `"extend-video"`, or `"reference-to-video"`. When you omit `mode`, the AI SDK uses standard generation.398Do not mix AI SDK `mode` values. Each request supports exactly one of `"edit-video"`, `"extend-video"`, or `"reference-to-video"`. When you omit `mode`, the AI SDK uses standard generation.

398 399 

399See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame) for `last_frame` examples.400See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame) for `last_frame` examples and [Keyframes](/developers/model-capabilities/video/reference-to-video#keyframes) for mid-video pins.

400 401 

401## Customize Polling Behavior402## Customize Polling Behavior

402 403 

Details

2 2 

3# Image-to-Video3# Image-to-Video

4 4 

5Transform a still image into a video by providing a source image along with an optional prompt. The model animates the image content based on your instructions. On `grok-imagine-video-1.5`, image-to-video supports native 1080p. Sending `image` alone is image-to-video. Combining it with `last_frame` or `reference_images` is reference-to-video with that image as the pinned first frame. See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame).5Transform a still image into a video by providing a source image along with an optional prompt. The model animates the image content based on your instructions. On `grok-imagine-video-1.5`, image-to-video supports native 1080p. Sending `image` alone is image-to-video. Combining it with `last_frame`, `keyframes`, or `reference_images` is reference-to-video with that image as the pinned first frame. See [First & Last frame](/developers/model-capabilities/video/reference-to-video#first--last-frame) and [Keyframes](/developers/model-capabilities/video/reference-to-video#keyframes).

6 6 

7You can provide the source image as:7You can provide the source image as:

8 8 

Details

2 2 

3# Reference-to-Video3# Reference-to-Video

4 4 

5Provide reference images, a preset voice, or both to guide the generated video. Images incorporate specific people, objects, clothing, or other visual elements without locking the first frame (unlike [image-to-video](/developers/model-capabilities/video/image-to-video)). This is useful for virtual try-on, product placement, character-consistent storytelling, and voice identity. On `grok-imagine-video-1.5`, you can also pick the voice your subject speaks in (see [Reference audio](#reference-audio)), and pin the exact first or last frame (see [First & Last frame](#first--last-frame)).5Provide reference images, a preset voice, or both to guide the generated video. Images incorporate specific people, objects, clothing, or other visual elements without locking the first frame (unlike [image-to-video](/developers/model-capabilities/video/image-to-video)). This is useful for virtual try-on, product placement, character-consistent storytelling, and voice identity. On `grok-imagine-video-1.5`, you can also pick the voice your subject speaks in (see [Reference audio](#reference-audio)), pin the exact first or last frame (see [First & Last frame](#first--last-frame)), and pin frames at chosen moments inside the clip (see [Keyframes](#keyframes)).

6 6 

7Each reference image can be provided as a public HTTPS URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.7Each reference image can be provided as a public HTTPS URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.

8 8 


343done343done

344```344```

345 345 

346## Keyframes

347 

348On `grok-imagine-video-1.5`, `keyframes` pins images at chosen moments inside the clip. Each entry pairs an `image` with a `timestamp_s`, and the video passes through that exact image at that time. Use it to storyboard a shot: the model generates the motion between your anchors rather than inventing the whole clip from one frame.

349 

350```json

351"keyframes": [

352 {"image": {"url": "<KEYFRAME_URL_1>"}, "timestamp_s": 2.0},

353 {"image": {"url": "<KEYFRAME_URL_2>"}, "timestamp_s": 4.0}

354]

355```

356 

357Keyframes cover the interior of the clip; the endpoints keep their own fields. Pin the opening with `image` and the closing with `last_frame`, and combine any of the three with `reference_images` or `reference_audios`. `prompt` is optional whenever a frame is pinned.

358 

359| Constraint | Detail |

360|------------|--------|

361| Count | At most 4 keyframes per request. |

362| Timing | `timestamp_s` must fall strictly inside the clip: greater than 0 and less than `duration`. Use `image` / `last_frame` for the endpoints. |

363| Spacing | Anchors snap to a 1/3-second grid. Two keyframes that round to the same slot are rejected, so keep them at least 1/3 s apart. |

364| Inputs | Each `image` accepts the same URL, data-URI, and `file_id` shapes as [image-to-video](/developers/model-capabilities/video/image-to-video). |

365 

366The Python SDK and Vercel AI SDK do not yet expose a dedicated `keyframes` parameter; send it on the REST body. Classic `grok-imagine-video` rejects `keyframes`, and keyframes cannot be combined with video editing.

367 

368```python customLanguage="pythonRequests"

369import os

370import time

371import requests

372 

373headers = {

374 "Content-Type": "application/json",

375 "Authorization": f"Bearer {os.environ['XAI_API_KEY']}",

376}

377 

378response = requests.post(

379 "https://api.x.ai/v1/videos/generations",

380 headers=headers,

381 json={

382 "model": "grok-imagine-video-1.5",

383 "prompt": "A slow tracking shot through the workshop: the sketch on the bench becomes a clay model, then the finished bronze in the window.",

384 "image": {"url": "<FIRST_FRAME_URL>"},

385 "keyframes": [

386 {"image": {"url": "<KEYFRAME_URL_1>"}, "timestamp_s": 3.0},

387 {"image": {"url": "<KEYFRAME_URL_2>"}, "timestamp_s": 6.0},

388 ],

389 "last_frame": {"url": "<LAST_FRAME_URL>"},

390 "duration": 8,

391 "aspect_ratio": "16:9",

392 "resolution": "720p",

393 },

394)

395 

396request_id = response.json()["request_id"]

397 

398while True:

399 result = requests.get(

400 f"https://api.x.ai/v1/videos/{request_id}",

401 headers={"Authorization": headers["Authorization"]},

402 )

403 data = result.json()

404 if data["status"] == "done":

405 print(data["video"]["url"])

406 break

407 elif data["status"] == "expired":

408 print("Request expired")

409 break

410 time.sleep(5)

411```

412 

413```bash

414REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \

415 -H "Content-Type: application/json" \

416 -H "Authorization: Bearer $XAI_API_KEY" \

417 -d '{

418 "model": "grok-imagine-video-1.5",

419 "prompt": "A slow tracking shot through the workshop: the sketch on the bench becomes a clay model, then the finished bronze in the window.",

420 "image": {"url": "<FIRST_FRAME_URL>"},

421 "keyframes": [

422 {"image": {"url": "<KEYFRAME_URL_1>"}, "timestamp_s": 3.0},

423 {"image": {"url": "<KEYFRAME_URL_2>"}, "timestamp_s": 6.0}

424 ],

425 "last_frame": {"url": "<LAST_FRAME_URL>"},

426 "duration": 8,

427 "aspect_ratio": "16:9",

428 "resolution": "720p"

429 }' | jq -r '.request_id')

430 

431while true; do

432 RESULT=$(curl -s https://api.x.ai/v1/videos/$REQUEST_ID \

433 -H "Authorization: Bearer $XAI_API_KEY")

434 STATUS=$(echo "$RESULT" | jq -r '.status')

435 if [ "$STATUS" = "done" ]; then

436 echo "$RESULT" | jq -r '.video.url'

437 break

438 elif [ "$STATUS" = "failed" ] || [ "$STATUS" = "expired" ]; then

439 echo "Request $STATUS"; echo "$RESULT" | jq .

440 break

441 fi

442 sleep 5

443done

444```

445 

346## Related446## Related

347 447 

348* [Video Generation](/developers/model-capabilities/video/generation) — Generate videos from text prompts448* [Video Generation](/developers/model-capabilities/video/generation) — Generate videos from text prompts

rate-limits.md +1 −1

Details

41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

44| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |45| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |48| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

49 49 

Details

24 Also accepts \`image\_url\` for compatibility.24 Also accepts \`image\_url\` for compatibility.

25 Required when \`file\_id\` is not set.25 Required when \`file\_id\` is not set.

26 26 

27* `keyframes` (array\<object>) — Optional mid-video keyframe anchors, strictly between the endpoint

28 pins (\`image\` as the first frame, \`last\_frame\` as the last). Each

29 entry pins an image to appear literally at its timestamp. Only

30 supported by select video models; at most 4 entries.

31 

32 * `image` (object, required) — Image input for generation and editing requests.

33 Accepts a public URL, a base64-encoded data URL, or a file\_id from the xAI Files API.

34 

35 * `file_id` (string | null) — File ID from the xAI Files API. Mutually exclusive with \`url\`.

36 The file must be an image (JPEG, PNG, or WebP) and fully uploaded.

37 

38 * `url` (string) — Public URL or base64-encoded data URL of the image (JPEG, PNG, or WebP).

39 Also accepts \`image\_url\` for compatibility.

40 Required when \`file\_id\` is not set.

41 

42 * `timestamp_s` (number, required) — Anchor time in seconds, strictly inside the clip (0 \< t \< duration).

43 

27* `model` (string | null) — Model to be used.44* `model` (string | null) — Model to be used.

28 45 

29* `output` (object)46* `output` (object)

Details

33 33 

34This includes **every tool call attempt**, even if some fail.34This includes **every tool call attempt**, even if some fail.

35 35 

36### `server_side_tool_usage` - Successful Calls (Billable)36### `server_side_tool_usage` - Successful Calls

37 37 

38```pythonWithoutSDK38```pythonWithoutSDK

39response.server_side_tool_usage39response.server_side_tool_usage

40```40```

41 41 

42Returns a map of successfully executed tools and their invocation counts. This represents only the tool calls that returned meaningful responses and **determines your billing**.42Returns a map of successfully executed tools and their invocation counts. This represents only the tool calls that returned meaningful responses and **determines your billing** for per-call tools; X Search is billed on the [item counts below](#x-search-item-counts) instead.

43 43 

44```output44```output

45{'SERVER_SIDE_TOOL_X_SEARCH': 3, 'SERVER_SIDE_TOOL_WEB_SEARCH': 2}45{'SERVER_SIDE_TOOL_X_SEARCH': 3, 'SERVER_SIDE_TOOL_WEB_SEARCH': 2}

46```46```

47 47 

48### X Search item counts

49 

50As of September 21, 2026, X Search is billed per post and per user profile fetched rather than per call. The Responses API reports these counts as `usage.server_side_tool_usage_details.x_posts_fetched` and `x_users_fetched`. See [X Search usage counts](/developers/tools/x-search#usage-counts) for what each field includes.

51 

48## Tool Call Function Names vs Usage Categories52## Tool Call Function Names vs Usage Categories

49 53 

50In xAI SDK chat responses, the function names in `tool_calls` represent the precise name of the tool invoked, while the entries in `server_side_tool_usage` provide a high-level categorization that aligns with the original tool passed in the `tools` array. In the Responses API, Web Search activity is represented as `web_search_call` output items instead.54In xAI SDK chat responses, the function names in `tool_calls` represent the precise name of the tool invoked, while the entries in `server_side_tool_usage` provide a high-level categorization that aligns with the original tool passed in the `tools` array. In the Responses API, Web Search activity is represented as `web_search_call` output items instead.


70 74 

71The agentic system handles these failures gracefully, updating its trajectory and continuing with alternative approaches when needed.75The agentic system handles these failures gracefully, updating its trajectory and continuing with alternative approaches when needed.

72 76 

73**Billing Note**: Only successful tool executions (`server_side_tool_usage`) are billed. Failed attempts are not charged.77**Billing Note**: Only successful tool executions (`server_side_tool_usage`) are billed. Failed attempts are not charged. X Search is billed per post and per user profile fetched rather than per call; see [X Search item counts](#x-search-item-counts).

74 78 

75## Understanding Token Usage79## Understanding Token Usage

76 80