SpyBara
Go Premium

Documentation 2026-08-25 22:58 UTC to 2026-08-31 20:58 UTC

12 files changed +27 −28. View all changes and history on the product overview
2026
Mon 31 20:58 Tue 25 22:58 Fri 21 18:57 Thu 20 15:58 Wed 19 18:02 Tue 18 04:01 Thu 13 22:00 Wed 12 23:59 Tue 11 20:57 Sat 8 23:00 Fri 7 17:57 Thu 6 20:01 Mon 3 23:00 Sat 1 01:59
Details

201 201 

202| Parameter | Type | Default | Description |202| Parameter | Type | Default | Description |

203|-----------|------|---------|-------------|203|-----------|------|---------|-------------|

204| `sample_rate` | integer | `16000` | Audio sample rate in Hz. With `encoding=opus`: `8000`, `16000`, `24000`, or `48000` only. |204| `sample_rate` | integer | `16000` | Audio sample rate in Hz. Ignored for `encoding=opus` — Opus packets are sample-rate-agnostic. |

205| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |205| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |

206| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |206| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |

207| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |207| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |


254 254 

255### Opus Streaming255### Opus Streaming

256 256 

257Set `encoding=opus` to stream compressed audio instead of raw PCM — roughly 4 KB/s at 24 kHz versus 48 KB/s for PCM16, with no client-side resampling.257Set `encoding=opus` to stream compressed audio instead of raw PCM — roughly 4 KB/s versus 48 KB/s for PCM16 at 24 kHz, with no client-side resampling.

258 258 

259**How it works:**259**How it works:**

260 260 

261* Send **exactly one raw Opus packet per binary WebSocket frame**. Opus packets don't mark their own boundaries — the WebSocket framing does — so never concatenate packets or split one across frames.261* Send **exactly one raw Opus packet per binary WebSocket frame**. Opus packets don't mark their own boundaries — the WebSocket framing does — so never concatenate packets or split one across frames.

262* Encode mono audio at `8000`, `16000`, `24000`, or `48000` Hz. Other sample rates are rejected with `400`.262* Omit `sample_rate` — Opus packets don't carry one, and the parameter is ignored for this encoding. Encode at whatever rate your audio source produces.

263* `multichannel` is not supported with Opus — mono only.263* `multichannel` is not supported with Opus — mono only.

264* Send raw packets, not containers. To transcribe an Ogg-Opus or WebM file, use the [batch endpoint](#supported-audio-formats) instead — it auto-detects containers.264* Send raw packets, not containers. To transcribe an Ogg-Opus or WebM file, use the [batch endpoint](#supported-audio-formats) instead — it auto-detects containers.

265* If a frame can't be decoded, the server sends an `error` event and closes the session. Misframed packets often decode as noise rather than an error, so double-check that your encoder emits one whole packet per frame.265* If a frame can't be decoded, the server sends an `error` event and closes the session. Misframed packets often decode as noise rather than an error, so double-check that your encoder emits one whole packet per frame.


267**Example URL:**267**Example URL:**

268 268 

269```269```

270wss://api.x.ai/v1/stt?sample_rate=24000&encoding=opus&interim_results=true270wss://api.x.ai/v1/stt?encoding=opus&interim_results=true

271```271```

272 272 

273**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.273**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.


441* **Enable `interim_results`** for responsive UX — show transcription as the user speaks441* **Enable `interim_results`** for responsive UX — show transcription as the user speaks

442* **Use `language=en`** to enable text formatting — numbers and currencies are written in their standard form442* **Use `language=en`** to enable text formatting — numbers and currencies are written in their standard form

443* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency443* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency

444* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s at 24 kHz versus 48 KB/s for raw PCM. See [Opus Streaming](#opus-streaming)444* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s versus 48 KB/s for raw PCM at 24 kHz. See [Opus Streaming](#opus-streaming)

445* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend445* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend

446 446 

447## Error Handling447## Error Handling

Details

337 337 

338### Quality338### Quality

339 339 

340Control generation quality with the optional `quality` parameter. Allowed values are `low` and `medium`. When omitted, the default is `medium`. The parameter is only supported for `grok-imagine-image-2.0`.340Control generation quality with the optional `quality` parameter. Allowed values are `low`, `medium`, and `auto`. When omitted, the default is `auto`, which lets the service choose the quality for each request. Auto currently uses `low` for image generation and `medium` for [image editing](/developers/model-capabilities/images/editing). Images are billed at the quality they are served at (see [Pricing](/developers/pricing)). Pass `low` or `medium` to pin a specific quality. The parameter is only supported for `grok-imagine-image-2.0`.

341 341 

342```python customLanguage="pythonXAI"342```python customLanguage="pythonXAI"

343import xai_sdk343import xai_sdk

Details

2 2 

3# Multi-Image Editing3# Multi-Image Editing

4 4 

5Use up to three source images for a single image edit. You can specify images in the order they are sent in the request. By default, the output aspect ratio follows the first input image. You can override this by setting the `aspect_ratio` parameter to a specific ratio, such as `"1:1"` or `"16:9"`.5Use up to five source images for a single image edit. You can specify images in the order they are sent in the request. By default, the output aspect ratio follows the first input image. You can override this by setting the `aspect_ratio` parameter to a specific ratio, such as `"1:1"` or `"16:9"`.

6 6 

7Each source image can be a public URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.7Each source image can be a public URL, a base64-encoded data URI, or a `file_id` from the [Files API](/developers/files) — and you can mix kinds within a single request. See [Imagine → Files API Integration](/developers/model-capabilities/imagine/files/inputs) for `file_id` details and examples.

8 8 

Details

2 2 

3# Imagine Overview3# Imagine Overview

4 4 

5The Imagine API lets you generate and edit images and videos with Grok Imagine models. Use it for image generation, image editing with up to 3 reference images, video generation from text or still images, video editing, and more.5The Imagine API lets you generate and edit images and videos with Grok Imagine models. Use it for image generation, image editing with up to 5 reference images, video generation from text or still images, video editing, and more.

6 6 

7## Pricing7## Pricing

8 8 


65 65 

66## Image Editing66## Image Editing

67 67 

68Edit a source image with natural language. Provide a public image URL or base64-encoded data URI, then describe the change you want Grok Imagine to apply. Multi-image editing supports up to 3 source images in a single request for combining subjects, transferring styles, and composing scenes.68Edit a source image with natural language. Provide a public image URL or base64-encoded data URI, then describe the change you want Grok Imagine to apply. Multi-image editing supports up to 5 source images in a single request for combining subjects, transferring styles, and composing scenes.

69 69 

70```python customLanguage="pythonXAI"70```python customLanguage="pythonXAI"

71import base6471import base64


190 190 

191Beyond the top use cases above, the Imagine API supports several additional workflows:191Beyond the top use cases above, the Imagine API supports several additional workflows:

192 192 

193* **[Multi-Image Editing](/developers/model-capabilities/images/multi-image-editing)** — Combine up to 3 source images in a single edit for compositing subjects, transferring styles, and building scenes from multiple references.193* **[Multi-Image Editing](/developers/model-capabilities/images/multi-image-editing)** — Combine up to 5 source images in a single edit for compositing subjects, transferring styles, and building scenes from multiple references.

194* **[Video Generation](/developers/model-capabilities/video/generation)** — Generate videos from text prompts with configurable duration (up to 15s), aspect ratio, and resolution.194* **[Video Generation](/developers/model-capabilities/video/generation)** — Generate videos from text prompts with configurable duration (up to 15s), aspect ratio, and resolution.

195* **[Video Editing](/developers/model-capabilities/video/editing)** — Modify an existing video with a text prompt while preserving the rest of the scene.195* **[Video Editing](/developers/model-capabilities/video/editing)** — Modify an existing video with a text prompt while preserving the rest of the scene.

196* **[Reference-to-Video](/developers/model-capabilities/video/reference-to-video)** — Guide a generated video with one or more reference images that influence the output without forcing the first frame.196* **[Reference-to-Video](/developers/model-capabilities/video/reference-to-video)** — Guide a generated video with one or more reference images that influence the output without forcing the first frame.

Details

77 messages: [77 messages: [

78 {78 {

79 role: "system",79 role: "system",

80 content: "You are Grok, a helpful and maximally truthful AI built by xAI."80 content: "You are Grok, a helpful and useful AI built by xAI."

81 },81 },

82 {82 {

83 role: "user",83 role: "user",


96const result = await generateText({96const result = await generateText({

97 model: xai('grok-4.6'),97 model: xai('grok-4.6'),

98 system:98 system:

99 "You are Grok, a helpful and maximally truthful AI built by xAI.",99 "You are Grok, a helpful and useful AI built by xAI.",

100 prompt: 'Explain how neural networks learn in two sentences.',100 prompt: 'Explain how neural networks learn in two sentences.',

101});101});

102 102 


112 "messages": [112 "messages": [

113 {113 {

114 "role": "system",114 "role": "system",

115 "content": "You are Grok, a helpful and maximally truthful AI built by xAI."115 "content": "You are Grok, a helpful and useful AI built by xAI."

116 },116 },

117 {117 {

118 "role": "user",118 "role": "user",

Details

28 28 

29chat = client.chat.create(model="grok-4.6")29chat = client.chat.create(model="grok-4.6")

30chat.append(30chat.append(

31 system("You are Grok, a helpful and maximally truthful AI built by xAI."),31 system("You are Grok, a helpful and useful AI built by xAI."),

32)32)

33chat.append(33chat.append(

34 user("Explain how neural networks learn in two sentences.")34 user("Explain how neural networks learn in two sentences.")


55stream = client.chat.completions.create(55stream = client.chat.completions.create(

56 model="grok-4.6",56 model="grok-4.6",

57 messages=[57 messages=[

58 {"role": "system", "content": "You are Grok, a helpful and maximally truthful AI built by xAI."},58 {"role": "system", "content": "You are Grok, a helpful and useful AI built by xAI."},

59 {"role": "user", "content": "Explain how neural networks learn in two sentences."},59 {"role": "user", "content": "Explain how neural networks learn in two sentences."},

60 ],60 ],

61 stream=True # Set streaming here61 stream=True # Set streaming here


76const stream = await openai.chat.completions.create({76const stream = await openai.chat.completions.create({

77 model: "grok-4.6",77 model: "grok-4.6",

78 messages: [78 messages: [

79 { role: "system", content: "You are Grok, a helpful and maximally truthful AI built by xAI." },79 { role: "system", content: "You are Grok, a helpful and useful AI built by xAI." },

80 {80 {

81 role: "user",81 role: "user",

82 content: "Explain how neural networks learn in two sentences.",82 content: "Explain how neural networks learn in two sentences.",


97const result = streamText({97const result = streamText({

98 model: xai.responses('grok-4.6'),98 model: xai.responses('grok-4.6'),

99 system:99 system:

100 "You are Grok, a helpful and maximally truthful AI built by xAI.",100 "You are Grok, a helpful and useful AI built by xAI.",

101 prompt: 'Explain how neural networks learn in two sentences.',101 prompt: 'Explain how neural networks learn in two sentences.',

102});102});

103 103 


115 "messages": [115 "messages": [

116 {116 {

117 "role": "system",117 "role": "system",

118 "content": "You are Grok, a helpful and maximally truthful AI built by xAI."118 "content": "You are Grok, a helpful and useful AI built by xAI."

119 },119 },

120 {120 {

121 "role": "user",121 "role": "user",

rate-limits.md +2 −2

Details

40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |43| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49### What counts toward TPM49### What counts toward TPM

50 50 

Details

348 348 

349* `object` (string, required) — The object type of this resource. Always set to \`response\`.349* `object` (string, required) — The object type of this resource. Always set to \`response\`.

350 350 

351* `output` (array\<object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.351* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

352 352 

353* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.353* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

354 354 


386 386 

387 * `type` (string, required) — Type is always \`"function"\`.387 * `type` (string, required) — Type is always \`"function"\`.

388 388 

389* `tools` (array\<object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.389* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

390 390 

391* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.391* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

392 392 


736 736 

737* `object` (string, required) — The object type of this resource. Always set to \`response\`.737* `object` (string, required) — The object type of this resource. Always set to \`response\`.

738 738 

739* `output` (array\<object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.739* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

740 740 

741* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.741* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

742 742 


774 774 

775 * `type` (string, required) — Type is always \`"function"\`.775 * `type` (string, required) — Type is always \`"function"\`.

776 776 

777* `tools` (array\<object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.777* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

778 778 

779* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.779* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

780 780 

Details

146 146 

147### Query Parameters147### Query Parameters

148 148 

149* `sample_rate` (integer, optional, default: 16000) — Audio sample rate in Hz. Supported values: \`8000\`, \`16000\`, \`22050\`, \`24000\`, \`44100\`, \`48000\`. With \`encoding=opus\`, only \`8000\`, \`16000\`, \`24000\`, and \`48000\` are supported.149* `sample_rate` (integer, optional, default: 16000) — Audio sample rate in Hz. Supported values: \`8000\`, \`16000\`, \`22050\`, \`24000\`, \`44100\`, \`48000\`. Ignored with \`encoding=opus\` — Opus packets are sample-rate-agnostic.

150 150 

151* `encoding` (string, optional, default: pcm) — Audio encoding format. \`pcm\` — signed 16-bit little-endian (2 bytes/sample). \`mulaw\` — G.711 µ-law (1 byte/sample). \`alaw\` — G.711 A-law (1 byte/sample). \`opus\` — raw Opus packets, one packet per binary WebSocket frame, mono only.151* `encoding` (string, optional, default: pcm) — Audio encoding format. \`pcm\` — signed 16-bit little-endian (2 bytes/sample). \`mulaw\` — G.711 µ-law (1 byte/sample). \`alaw\` — G.711 A-law (1 byte/sample). \`opus\` — raw Opus packets, one packet per binary WebSocket frame, mono only.

152 152 

Details

1064 1064 

1065### Query Parameters1065### Query Parameters

1066 1066 

1067* `sample_rate` (integer, optional, default: 16000) — Audio sample rate in Hz. Supported values: \`8000\`, \`16000\`, \`22050\`, \`24000\`, \`44100\`, \`48000\`. With \`encoding=opus\`, only \`8000\`, \`16000\`, \`24000\`, and \`48000\` are supported.1067* `sample_rate` (integer, optional, default: 16000) — Audio sample rate in Hz. Supported values: \`8000\`, \`16000\`, \`22050\`, \`24000\`, \`44100\`, \`48000\`. Ignored with \`encoding=opus\` — Opus packets are sample-rate-agnostic.

1068 1068 

1069* `encoding` (string, optional, default: pcm) — Audio encoding format. \`pcm\` — signed 16-bit little-endian (2 bytes/sample). \`mulaw\` — G.711 µ-law (1 byte/sample). \`alaw\` — G.711 A-law (1 byte/sample). \`opus\` — raw Opus packets, one packet per binary WebSocket frame, mono only.1069* `encoding` (string, optional, default: pcm) — Audio encoding format. \`pcm\` — signed 16-bit little-endian (2 bytes/sample). \`mulaw\` — G.711 µ-law (1 byte/sample). \`alaw\` — G.711 A-law (1 byte/sample). \`opus\` — raw Opus packets, one packet per binary WebSocket frame, mono only.

1070 1070 

Details

47 * `email` (string) — User's email. May not always populated.47 * `email` (string) — User's email. May not always populated.

48 48 

49 * `profileImage` (string) — The key of the profile image under which it can be fetched from our assets server.49 * `profileImage` (string) — The key of the profile image under which it can be fetched from our assets server.

50 TODO(pohlen): This should be the profile picture URL.

51 50 

52 * `givenName` (string) — User's given name.51 * `givenName` (string) — User's given name.

53 52 

Details

292 "content": [292 "content": [

293 {293 {

294 "type": "output_text",294 "type": "output_text",

295 "text": "**xAI is an artificial intelligence company founded by Elon Musk in March 2023.** Its stated mission is to \"understand the universe\" by building advanced AI systems that accelerate human scientific discovery.[[1]](https://x.ai/company)\n\n### Key Details\n- **Flagship product**: Grok, a family of frontier AI models focused on reasoning, code, voice, image generation, and video. These are trained on massive infrastructure, including what the company describes as the world's largest supercluster (Colossus). Grok powers chatbots, APIs, and multimodal tools available via a unified API.[[2]](https://x.ai/)\n- **Current status (as of mid-2026)**: xAI operates as a subsidiary of SpaceX following an acquisition in February 2026. It is also connected to the X social platform (formerly Twitter), which xAI effectively became the parent of in 2025. The company has expanded into data centers and enterprise AI offerings (e.g., integrations with Amazon Bedrock and Databricks).[[3]](https://en.wikipedia.org/wiki/XAI_(company))\n- **Headquarters and team**: Based in the Stanford Research Park in Palo Alto, California. It was initially founded with a team of AI researchers and is led by Elon Musk as CEO.\n\nxAI positions itself as building maximally truth-seeking AI, distinct from other labs in its approach. Its official website (x.ai) highlights developer tools, API access, and ongoing model releases. Note that there is an unrelated blockchain/gaming project called Xai (xai.games), but the primary reference to \"xAI\" in this context is Musk's AI venture.[[4]](https://xai.games/)\n\nFor the latest updates, check x.ai or @xai on X.",295 "text": "**xAI is an artificial intelligence company founded by Elon Musk in March 2023.** Its stated mission is to \"understand the universe\" by building advanced AI systems that accelerate human scientific discovery.[[1]](https://x.ai/company)\n\n### Key Details\n- **Flagship product**: Grok, a family of frontier AI models focused on reasoning, code, voice, image generation, and video. These are trained on massive infrastructure, including what the company describes as the world's largest supercluster (Colossus). Grok powers chatbots, APIs, and multimodal tools available via a unified API.[[2]](https://x.ai/)\n- **Current status (as of mid-2026)**: xAI operates as a subsidiary of SpaceX following an acquisition in February 2026. It is also connected to the X social platform (formerly Twitter), which xAI effectively became the parent of in 2025. The company has expanded into data centers and enterprise AI offerings (e.g., integrations with Amazon Bedrock and Databricks).[[3]](https://en.wikipedia.org/wiki/XAI_(company))\n- **Headquarters and team**: Based in the Stanford Research Park in Palo Alto, California. It was initially founded with a team of AI researchers and is led by Elon Musk as CEO.\n\nxAI positions itself as building useful AI, distinct from other labs in its approach. Its official website (x.ai) highlights developer tools, API access, and ongoing model releases. Note that there is an unrelated blockchain/gaming project called Xai (xai.games), but the primary reference to \"xAI\" in this context is Musk's AI venture.[[4]](https://xai.games/)\n\nFor the latest updates, check x.ai or @xai on X.",

296 "logprobs": [],296 "logprobs": [],

297 "annotations": [297 "annotations": [

298 {298 {