101 101
102| Parameter | Type | Required | Description |102| Parameter | Type | Required | Description |
103|-----------|------|----------|-------------|103|-----------|------|----------|-------------|
104104| `text` | string | ✓ | The text to convert to speech. Maximum **15,000 characters**. Supports [speech tags](#speech-tags). || `text` | string | ✓ | The text to convert to speech. Maximum **60,000 characters**. Supports [speech tags](#speech-tags). |
105| `voice_id` | string | | Voice to use for synthesis. Defaults to `eve`. See [Voices](#voices). |105| `voice_id` | string | | Voice to use for synthesis. Defaults to `eve`. See [Voices](#voices). |
106| `language` | string | ✓ | BCP-47 language code (e.g. `en`, `zh`, `pt-BR`) or `auto` for automatic language detection. See [Supported Languages](#supported-languages). |106| `language` | string | ✓ | BCP-47 language code (e.g. `en`, `zh`, `pt-BR`) or `auto` for automatic language detection. See [Supported Languages](#supported-languages). |
107| `output_format` | object | | Output format configuration. Defaults to MP3 at 24 kHz / 128 kbps. See [Output Formats](#output-formats). |107| `output_format` | object | | Output format configuration. Defaults to MP3 at 24 kHz / 128 kbps. See [Output Formats](#output-formats). |
729| Key characters | letters, digits, apostrophes, spaces | `replace key "C++" may not contain punctuation or symbols` |729| Key characters | letters, digits, apostrophes, spaces | `replace key "C++" may not contain punctuation or symbols` |
730| Keys non-blank | — | `replace keys must not be blank` |730| Keys non-blank | — | `replace keys must not be blank` |
731| Keys distinct as phrases | compared case- and whitespace-insensitively | `replace keys "ACME" and "Acme" are the same phrase; keep one` |731| Keys distinct as phrases | compared case- and whitespace-insensitively | `replace keys "ACME" and "Acme" are the same phrase; keep one` |
732732| Text after substitution | 60,000 characters | `` `replace` expands the text to … characters `` || Text after substitution | 240,000 characters | `` `replace` expands the text to … characters `` |
733 733
734734The last row bounds the **rewritten** text — watch it when a short key maps to a long value across a large body of text. Your 15,000-character input cap and your billing still count what you sent.The last row bounds the **rewritten** text — watch it when a short key maps to a long value across a large body of text. Your 60,000-character input cap and your billing still count what you sent.
735 735
736The same `replace` map is available on the [Speech to Speech API](/developers/model-capabilities/audio/speech-to-speech#pronunciation-replacements) and on the [WebSocket endpoint](#session-configuration).736The same `replace` map is available on the [Speech to Speech API](/developers/model-capabilities/audio/speech-to-speech#pronunciation-replacements) and on the [WebSocket endpoint](#session-configuration).
737 737
806* **Use natural punctuation.** Commas, periods, and question marks guide pacing and intonation. `"Wait, really?"` sounds more natural than `"Wait really"`.806* **Use natural punctuation.** Commas, periods, and question marks guide pacing and intonation. `"Wait, really?"` sounds more natural than `"Wait really"`.
807* **Add emotional context.** Exclamation marks and question marks influence delivery - `"That's amazing!"` sounds enthusiastic while `"That's amazing."` is matter-of-fact.807* **Add emotional context.** Exclamation marks and question marks influence delivery - `"That's amazing!"` sounds enthusiastic while `"That's amazing."` is matter-of-fact.
808* **Break long content into paragraphs.** Paragraph breaks create natural pauses and help the model maintain consistent quality across longer text.808* **Break long content into paragraphs.** Paragraph breaks create natural pauses and help the model maintain consistent quality across longer text.
809809* **Keep unary requests under 15,000 characters.** For longer content, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit, or split into logical segments (by paragraph or sentence) and concatenate the audio output.* **Keep unary requests under 60,000 characters.** For longer content, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit, or split into logical segments (by paragraph or sentence) and concatenate the audio output.
810 810
811### Integrating with AI coding assistants811### Integrating with AI coding assistants
812 812
916| Status | Meaning | Action |916| Status | Meaning | Action |
917|--------|---------|--------|917|--------|---------|--------|
918| `200` | Success | Audio bytes in the response body |918| `200` | Success | Audio bytes in the response body |
919919| `400` | Bad request | Check: text is non-empty, under 15,000 chars; codec and sample rate are valid || `400` | Bad request | Check: text is non-empty, under 60,000 chars; codec and sample rate are valid |
920| `401` | Unauthorized | API key is missing or invalid |920| `401` | Unauthorized | API key is missing or invalid |
921| `404` | Not found | Unknown `voice_id` — verify via `GET /v1/tts/voices` (built-in) or `GET /v1/custom-voices` (custom) |921| `404` | Not found | Unknown `voice_id` — verify via `GET /v1/tts/voices` (built-in) or `GET /v1/custom-voices` (custom) |
922| `429` | Rate limited | Back off and retry with exponential delay |922| `429` | Rate limited | Back off and retry with exponential delay |
1009 1009
1010| | Unary & server-streamed (`POST /v1/tts`) | Bidirectional WebSocket (`wss://api.x.ai/v1/tts`) |1010| | Unary & server-streamed (`POST /v1/tts`) | Bidirectional WebSocket (`wss://api.x.ai/v1/tts`) |
1011|---|---:|---|1011|---|---:|---|
10121012| **Max text length** | 15,000 characters per request | No limit — individual `text.delta` messages capped at 15,000 characters each || **Max text length** | 60,000 characters per request | No limit — individual `text.delta` messages capped at 60,000 characters each |
1013| **Request timeout** | 15 minutes | No timeout (connection stays open) |1013| **Request timeout** | 15 minutes | No timeout (connection stays open) |
1014| **Concurrent sessions** | — | 50 per team |1014| **Concurrent sessions** | — | 50 per team |
1015| **[`replace`](#map-limits) map** | 200 entries; keys ≤ 100 and values ≤ 128 characters | Same, per `session.update` |1015| **[`replace`](#map-limits) map** | 200 entries; keys ≤ 100 and values ≤ 128 characters | Same, per `session.update` |
10161016| **Text after `replace`** | 60,000 characters | 60,000 characters per utterance || **Text after `replace`** | 240,000 characters | 240,000 characters per utterance |
1017 1017
10181018For content exceeding 15,000 characters, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit.For content exceeding 60,000 characters, use the [bidirectional WebSocket endpoint](#streaming-tts-websocket) which has no text length limit.
1019 1019
1020## Streaming TTS (WebSocket)1020## Streaming TTS (WebSocket)
1021 1021
1063 1063
1064| Event | Description |1064| Event | Description |
1065|-------|-------------|1065|-------|-------------|
10661066| `text.delta` | A chunk of text to synthesize. Individual deltas are capped at **15,000 characters**. || `text.delta` | A chunk of text to synthesize. Individual deltas are capped at **60,000 characters**. |
1067| `text.done` | Signals the end of the current utterance. The server will finish generating audio and send `audio.done`. |1067| `text.done` | Signals the end of the current utterance. The server will finish generating audio and send `audio.done`. |
1068| `text.clear` | Cancel the current utterance. The server stops generating audio, discards any buffered data, and responds with `audio.clear`. |1068| `text.clear` | Cancel the current utterance. The server stops generating audio, discards any buffered data, and responds with `audio.clear`. |
1069| `session.update` | Set or change the [`replace`](#session-configuration) map for the session. Accepted at any point; it takes effect on the next utterance to begin. |1069| `session.update` | Set or change the [`replace`](#session-configuration) map for the session. Accepted at any point; it takes effect on the next utterance to begin. |
1107 1107
1108Matching runs across `text.delta` boundaries, so a phrase split over two messages still matches.1108Matching runs across `text.delta` boundaries, so a phrase split over two messages still matches.
1109 1109
11101110The [same map limits](#map-limits) apply. A map that fails validation is answered with an `error` frame, leaves the map in effect unchanged, and keeps the connection open — but a map that expands a turn past 60,000 characters ends the session, since by then the oversized text has already been accepted.The [same map limits](#map-limits) apply. A map that fails validation is answered with an `error` frame, leaves the map in effect unchanged, and keeps the connection open — but a map that expands a turn past 240,000 characters ends the session, since by then the oversized text has already been accepted.
1111 1111
1112### Multi-Utterance Sessions1112### Multi-Utterance Sessions
1113 1113
1435| Property | Value |1435| Property | Value |
1436|----------|-------|1436|----------|-------|
1437| **Total text length** | No limit — send as many `text.delta` messages as needed |1437| **Total text length** | No limit — send as many `text.delta` messages as needed |
14381438| **Delta size** | Individual `text.delta` messages capped at 15,000 characters || **Delta size** | Individual `text.delta` messages capped at 60,000 characters |
1439| **Concurrent sessions** | 50 per team |1439| **Concurrent sessions** | 50 per team |
1440| **Session permit TTL** | 600 seconds |1440| **Session permit TTL** | 600 seconds |
14411441| **[`replace`](#map-limits) expansion** | An utterance whose text exceeds 60,000 characters after substitution ends the session || **[`replace`](#map-limits) expansion** | An utterance whose text exceeds 240,000 characters after substitution ends the session |
1442| **Moderation** | Runs asynchronously on accumulated text after audio is sent (fail-open) |1442| **Moderation** | Runs asynchronously on accumulated text after audio is sent (fail-open) |
1443| **Billing** | Recorded per session based on total input characters |1443| **Billing** | Recorded per session based on total input characters |
1444 1444