SpyBara
Go Premium

Documentation 2026-09-15 17:00 UTC to 2026-09-16 19:01 UTC

8 files changed +22 −98. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03
Details

101 101 

102| Parameter | Type | Required | Description |102| Parameter | Type | Required | Description |

103|-----------|------|----------|-------------|103|-----------|------|----------|-------------|

104| `text` | string | ✓ | The text to convert to speech. Maximum **15,000 characters**. Supports [speech tags](#speech-tags). |104| `text` | string | ✓ | The text to convert to speech. Maximum **60,000 characters**. Supports [speech tags](#speech-tags). |

105| `voice_id` | string | | Voice to use for synthesis. Defaults to `eve`. See [Voices](#voices). |105| `voice_id` | string | | Voice to use for synthesis. Defaults to `eve`. See [Voices](#voices). |

106| `language` | string | ✓ | BCP-47 language code (e.g. `en`, `zh`, `pt-BR`) or `auto` for automatic language detection. See [Supported Languages](#supported-languages). |106| `language` | string | ✓ | BCP-47 language code (e.g. `en`, `zh`, `pt-BR`) or `auto` for automatic language detection. See [Supported Languages](#supported-languages). |

107| `output_format` | object | | Output format configuration. Defaults to MP3 at 24 kHz / 128 kbps. See [Output Formats](#output-formats). |107| `output_format` | object | | Output format configuration. Defaults to MP3 at 24 kHz / 128 kbps. See [Output Formats](#output-formats). |


729| Key characters | letters, digits, apostrophes, spaces | `replace key "C++" may not contain punctuation or symbols` |729| Key characters | letters, digits, apostrophes, spaces | `replace key "C++" may not contain punctuation or symbols` |

730| Keys non-blank | — | `replace keys must not be blank` |730| Keys non-blank | — | `replace keys must not be blank` |

731| Keys distinct as phrases | compared case- and whitespace-insensitively | `replace keys "ACME" and "Acme" are the same phrase; keep one` |731| Keys distinct as phrases | compared case- and whitespace-insensitively | `replace keys "ACME" and "Acme" are the same phrase; keep one` |

732| Text after substitution | 60,000 characters | `` `replace` expands the text to … characters `` |732| Text after substitution | 240,000 characters | `` `replace` expands the text to … characters `` |

733 733 

734The last row bounds the **rewritten** text — watch it when a short key maps to a long value across a large body of text. Your 15,000-character input cap and your billing still count what you sent.734The last row bounds the **rewritten** text — watch it when a short key maps to a long value across a large body of text. Your 60,000-character input cap and your billing still count what you sent.

735 735 

736The same `replace` map is available on the [Speech to Speech API](/developers/model-capabilities/audio/speech-to-speech#pronunciation-replacements) and on the [WebSocket endpoint](#session-configuration).736The same `replace` map is available on the [Speech to Speech API](/developers/model-capabilities/audio/speech-to-speech#pronunciation-replacements) and on the [WebSocket endpoint](#session-configuration).

737 737 


806* **Use natural punctuation.** Commas, periods, and question marks guide pacing and intonation. `"Wait, really?"` sounds more natural than `"Wait really"`.806* **Use natural punctuation.** Commas, periods, and question marks guide pacing and intonation. `"Wait, really?"` sounds more natural than `"Wait really"`.

807* **Add emotional context.** Exclamation marks and question marks influence delivery - `"That's amazing!"` sounds enthusiastic while `"That's amazing."` is matter-of-fact.807* **Add emotional context.** Exclamation marks and question marks influence delivery - `"That's amazing!"` sounds enthusiastic while `"That's amazing."` is matter-of-fact.

808* **Break long content into paragraphs.** Paragraph breaks create natural pauses and help the model maintain consistent quality across longer text.808* **Break long content into paragraphs.** Paragraph breaks create natural pauses and help the model maintain consistent quality across longer text.

809* **Keep unary requests under 15,000 characters.** For longer content, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit, or split into logical segments (by paragraph or sentence) and concatenate the audio output.809* **Keep unary requests under 60,000 characters.** For longer content, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit, or split into logical segments (by paragraph or sentence) and concatenate the audio output.

810 810 

811### Integrating with AI coding assistants811### Integrating with AI coding assistants

812 812 


916| Status | Meaning | Action |916| Status | Meaning | Action |

917|--------|---------|--------|917|--------|---------|--------|

918| `200` | Success | Audio bytes in the response body |918| `200` | Success | Audio bytes in the response body |

919| `400` | Bad request | Check: text is non-empty, under 15,000 chars; codec and sample rate are valid |919| `400` | Bad request | Check: text is non-empty, under 60,000 chars; codec and sample rate are valid |

920| `401` | Unauthorized | API key is missing or invalid |920| `401` | Unauthorized | API key is missing or invalid |

921| `404` | Not found | Unknown `voice_id` — verify via `GET /v1/tts/voices` (built-in) or `GET /v1/custom-voices` (custom) |921| `404` | Not found | Unknown `voice_id` — verify via `GET /v1/tts/voices` (built-in) or `GET /v1/custom-voices` (custom) |

922| `429` | Rate limited | Back off and retry with exponential delay |922| `429` | Rate limited | Back off and retry with exponential delay |


1009 1009 

1010| | Unary & server-streamed (`POST /v1/tts`) | Bidirectional WebSocket (`wss://api.x.ai/v1/tts`) |1010| | Unary & server-streamed (`POST /v1/tts`) | Bidirectional WebSocket (`wss://api.x.ai/v1/tts`) |

1011|---|---:|---|1011|---|---:|---|

1012| **Max text length** | 15,000 characters per request | No limit — individual `text.delta` messages capped at 15,000 characters each |1012| **Max text length** | 60,000 characters per request | No limit — individual `text.delta` messages capped at 60,000 characters each |

1013| **Request timeout** | 15 minutes | No timeout (connection stays open) |1013| **Request timeout** | 15 minutes | No timeout (connection stays open) |

1014| **Concurrent sessions** | — | 50 per team |1014| **Concurrent sessions** | — | 50 per team |

1015| **[`replace`](#map-limits) map** | 200 entries; keys ≤ 100 and values ≤ 128 characters | Same, per `session.update` |1015| **[`replace`](#map-limits) map** | 200 entries; keys ≤ 100 and values ≤ 128 characters | Same, per `session.update` |

1016| **Text after `replace`** | 60,000 characters | 60,000 characters per utterance |1016| **Text after `replace`** | 240,000 characters | 240,000 characters per utterance |

1017 1017 

1018For content exceeding 15,000 characters, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit.1018For content exceeding 60,000 characters, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit.

1019 1019 

1020## Streaming TTS (WebSocket)1020## Streaming TTS (WebSocket)

1021 1021 


1063 1063 

1064| Event | Description |1064| Event | Description |

1065|-------|-------------|1065|-------|-------------|

1066| `text.delta` | A chunk of text to synthesize. Individual deltas are capped at **15,000 characters**. |1066| `text.delta` | A chunk of text to synthesize. Individual deltas are capped at **60,000 characters**. |

1067| `text.done` | Signals the end of the current utterance. The server will finish generating audio and send `audio.done`. |1067| `text.done` | Signals the end of the current utterance. The server will finish generating audio and send `audio.done`. |

1068| `text.clear` | Cancel the current utterance. The server stops generating audio, discards any buffered data, and responds with `audio.clear`. |1068| `text.clear` | Cancel the current utterance. The server stops generating audio, discards any buffered data, and responds with `audio.clear`. |

1069| `session.update` | Set or change the [`replace`](#session-configuration) map for the session. Accepted at any point; it takes effect on the next utterance to begin. |1069| `session.update` | Set or change the [`replace`](#session-configuration) map for the session. Accepted at any point; it takes effect on the next utterance to begin. |


1107 1107 

1108Matching runs across `text.delta` boundaries, so a phrase split over two messages still matches.1108Matching runs across `text.delta` boundaries, so a phrase split over two messages still matches.

1109 1109 

1110The [same map limits](#map-limits) apply. A map that fails validation is answered with an `error` frame, leaves the map in effect unchanged, and keeps the connection open — but a map that expands a turn past 60,000 characters ends the session, since by then the oversized text has already been accepted.1110The [same map limits](#map-limits) apply. A map that fails validation is answered with an `error` frame, leaves the map in effect unchanged, and keeps the connection open — but a map that expands a turn past 240,000 characters ends the session, since by then the oversized text has already been accepted.

1111 1111 

1112### Multi-Utterance Sessions1112### Multi-Utterance Sessions

1113 1113 


1435| Property | Value |1435| Property | Value |

1436|----------|-------|1436|----------|-------|

1437| **Total text length** | No limit — send as many `text.delta` messages as needed |1437| **Total text length** | No limit — send as many `text.delta` messages as needed |

1438| **Delta size** | Individual `text.delta` messages capped at 15,000 characters |1438| **Delta size** | Individual `text.delta` messages capped at 60,000 characters |

1439| **Concurrent sessions** | 50 per team |1439| **Concurrent sessions** | 50 per team |

1440| **Session permit TTL** | 600 seconds |1440| **Session permit TTL** | 600 seconds |

1441| **[`replace`](#map-limits) expansion** | An utterance whose text exceeds 60,000 characters after substitution ends the session |1441| **[`replace`](#map-limits) expansion** | An utterance whose text exceeds 240,000 characters after substitution ends the session |

1442| **Moderation** | Runs asynchronously on accumulated text after audio is sent (fail-open) |1442| **Moderation** | Runs asynchronously on accumulated text after audio is sent (fail-open) |

1443| **Billing** | Recorded per session based on total input characters |1443| **Billing** | Recorded per session based on total input characters |

1444 1444 

rate-limits.md +3 −3

Details

40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |43| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49**Voice & Audio**49**Voice & Audio**

50 50 

Details

9* [Upload](/developers/rest-api-reference/files/upload)9* [Upload](/developers/rest-api-reference/files/upload)

10* [Manage](/developers/rest-api-reference/files/manage)10* [Manage](/developers/rest-api-reference/files/manage)

11* [Download](/developers/rest-api-reference/files/download)11* [Download](/developers/rest-api-reference/files/download)

12* [Public URLs](/developers/rest-api-reference/files/public-urls)

rest-api-reference/files/public-urls.md +0 −75 deleted

File Deleted View Diff

1#### Files API

2 

3# Public URLs

4 

5See the [Public URLs guide](/developers/files/public-urls) for expiry behaviour, idempotency, and end-to-end examples.

6 

7***

8 

9## POST /v1/files/\{file\_id}/public-url

10 

11Create a permanent, unauthenticated public URL for an existing file. The

12underlying file is unaffected and can still be fetched through the

13authenticated content endpoint. Use this when you want to share a stored

14asset (image, video, PDF) outside your API-keyed environment. Public URLs

15can be revoked at any time via \`POST /v1/files/\{file\_id}/public-url/revoke\`.

16 

17### Path Parameters

18 

19* `file_id` (string, required) — The file's \`id\`.

20 

21### Request Body

22 

23* `expires_after` (integer | null) — Seconds from now until the public URL expires. Must be between \`3600\` (1

24 hour) and \`2592000\` (30 days). Omit to inherit the file's expiry (if it

25 has one) or to make the URL valid indefinitely.

26 

27### Response Body

28 

29* `expires_at` (integer | null) — Unix timestamp (seconds) when the public URL expires. Present when

30 the public URL has an expiry, either from an explicit \`expires\_after\`

31 in the request or inherited from the file's TTL. Absent when the

32 public URL is valid indefinitely.

33 

34* `public_url` (string, required) — The full public URL.

35 

36\*\*Response example:\*\*

37 

38```json

39{

40 "public_url": "https://files-cdn.x.ai/ZsqeMtdcSYWPPHTQdxXDKQ/file_a128090d-f0c9-4873-bd84-e499777e7417.png",

41 "expires_at": 1755600000

42}

43```

44 

45***

46 

47## POST /v1/files/\{file\_id}/public-url/revoke

48 

49Revoke the active public URL for a file. The underlying file remains

50available through the authenticated content endpoint. Revoke is idempotent

51— calling it on a file without an active public URL returns

52\`revoked: false\` without an error.

53 

54### Path Parameters

55 

56* `file_id` (string, required) — The file's \`id\`.

57 

58### Response Body

59 

60* `id` (string, required) — The file ID whose public URL was revoked.

61 

62* `public_url` (string | null) — The full public URL that was revoked. Only present when \`revoked\` is \`true\`.

63 

64* `revoked` (boolean, required) — Whether a public URL was actually revoked. \`false\` if the file had no

65 active public URL (no-op).

66 

67\*\*Response example:\*\*

68 

69```json

70{

71 "id": "file_a128090d-f0c9-4873-bd84-e499777e7417",

72 "revoked": true,

73 "public_url": "https://files-cdn.x.ai/ZsqeMtdcSYWPPHTQdxXDKQ/file_a128090d-f0c9-4873-bd84-e499777e7417.png"

74}

75```

Details

303 303 

304 * `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.304 * `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

305 305 

306 * `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.306 * `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Higher efforts use more reasoning tokens for deeper thinking. The supported values and the default depend on the model.

307 307 

308 * `response_format` (object | object | object)308 * `response_format` (object | object | object)

309 309 

Details

38 across requests sharing a prompt prefix. Plumbed to \`x-grok-conv-id\`,38 across requests sharing a prompt prefix. Plumbed to \`x-grok-conv-id\`,

39 same as on \`/v1/responses\`.39 same as on \`/v1/responses\`.

40 40 

41* `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.41* `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Higher efforts use more reasoning tokens for deeper thinking. The supported values and the default depend on the model.

42 42 

43* `response_format` (object | object | object)43* `response_format` (object | object | object)

44 44 

Details

46 46 

47* `reasoning` (object)47* `reasoning` (object)

48 48 

49 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.49 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Higher efforts use more reasoning tokens for deeper thinking. The supported values and the default depend on the model.

50 50 

51 * `generate_summary` (string | null) — Only included for compatibility.51 * `generate_summary` (string | null) — Only included for compatibility.

52 52 


141 141 

142* `reasoning` (object)142* `reasoning` (object)

143 143 

144 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.144 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Higher efforts use more reasoning tokens for deeper thinking. The supported values and the default depend on the model.

145 145 

146 * `generate_summary` (string | null) — Only included for compatibility.146 * `generate_summary` (string | null) — Only included for compatibility.

147 147 


394 394 

395* `reasoning` (object)395* `reasoning` (object)

396 396 

397 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.397 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Higher efforts use more reasoning tokens for deeper thinking. The supported values and the default depend on the model.

398 398 

399 * `generate_summary` (string | null) — Only included for compatibility.399 * `generate_summary` (string | null) — Only included for compatibility.

400 400 

Details

393 393 

394### Request Body394### Request Body

395 395 

396* `text` (string, required) — The text to convert to speech. Maximum 15,000 characters. Supports inline speech tags for expressive output: \`\[pause]\`, \`\[long-pause]\`, \`\[hum-tune]\`, \`\[laugh]\`, \`\[chuckle]\`, \`\[giggle]\`, \`\[cry]\`, \`\[tsk]\`, \`\[tongue-click]\`, \`\[lip-smack]\`, \`\[breath]\`, \`\[inhale]\`, \`\[exhale]\`, \`\[sigh]\`. Also supports wrapping tags for style control: \`\<soft>\`, \`\<whisper>\`, \`\<loud>\`, \`\<build-intensity>\`, \`\<decrease-intensity>\`, \`\<higher-pitch>\`, \`\<lower-pitch>\`, \`\<slow>\`, \`\<fast>\`, \`\<sing-song>\`, \`\<singing>\`, \`\<emphasis>\`.396* `text` (string, required) — The text to convert to speech. Maximum 60,000 characters. Supports inline speech tags for expressive output: \`\[pause]\`, \`\[long-pause]\`, \`\[hum-tune]\`, \`\[laugh]\`, \`\[chuckle]\`, \`\[giggle]\`, \`\[cry]\`, \`\[tsk]\`, \`\[tongue-click]\`, \`\[lip-smack]\`, \`\[breath]\`, \`\[inhale]\`, \`\[exhale]\`, \`\[sigh]\`. Also supports wrapping tags for style control: \`\<soft>\`, \`\<whisper>\`, \`\<loud>\`, \`\<build-intensity>\`, \`\<decrease-intensity>\`, \`\<higher-pitch>\`, \`\<lower-pitch>\`, \`\<slow>\`, \`\<fast>\`, \`\<sing-song>\`, \`\<singing>\`, \`\<emphasis>\`.

397 397 

398* `voice_id` (string) — Voice identifier. Use a built-in voice from \`GET /v1/tts/voices\` (e.g. \`eve\`, \`ara\`) or a custom voice ID. Defaults to \`eve\` when omitted.398* `voice_id` (string) — Voice identifier. Use a built-in voice from \`GET /v1/tts/voices\` (e.g. \`eve\`, \`ara\`) or a custom voice ID. Defaults to \`eve\` when omitted.

399 399 


637 637 

638### Client Messages638### Client Messages

639 639 

640* `text.delta` — Send a chunk of text to be synthesized. Text is processed incrementally — audio generation begins as soon as enough text is buffered. Individual deltas are capped at 15,000 characters.640* `text.delta` — Send a chunk of text to be synthesized. Text is processed incrementally — audio generation begins as soon as enough text is buffered. Individual deltas are capped at 60,000 characters.

641 641 

642* `text.done` — Signal that all text for this utterance has been sent. The server will finish generating audio and send \`audio.done\`. After receiving \`audio.done\`, you can start a new utterance with another \`text.delta\`.642* `text.done` — Signal that all text for this utterance has been sent. The server will finish generating audio and send \`audio.done\`. After receiving \`audio.done\`, you can start a new utterance with another \`text.delta\`.

643 643