SpyBara
Go Premium

Documentation 2026-10-04 23:58 UTC to 2026-10-05 22:59 UTC

8 files changed +62 −39. View all changes and history on the product overview
2026
Sun 11 03:59 Sat 10 23:59 Fri 9 23:59 Thu 8 23:58 Wed 7 23:59 Tue 6 22:57 Mon 5 22:59 Sun 4 23:58 Sat 3 23:58 Fri 2 23:57 Thu 1 23:57
Details

309 309 

310* **`limit`**: Maximum number of files to return. If not specified, uses the server default of 100. Maximum is 100.310* **`limit`**: Maximum number of files to return. If not specified, uses the server default of 100. Maximum is 100.

311* **`order`**: Sort order. Either `"asc"` (ascending) or `"desc"` (descending). Defaults to `"desc"`.311* **`order`**: Sort order. Either `"asc"` (ascending) or `"desc"` (descending). Defaults to `"desc"`.

312* **`sort_by`**: Field to sort by. Options: `"created_at"`, `"filename"`, or `"size"`. Defaults to `"created_at"`.312* **`sort_by`**: Field to sort by. Options: `"created_at"`, `"filename"`, or `"size"`. Defaults to `"filename"`.

313* **`pagination_token`**: Pass the `pagination_token` returned by the previous response to fetch the next page. Omit it for the first page.313* **`pagination_token`**: Pass the `pagination_token` returned by the previous response to fetch the next page. Omit it for the first page.

314 314 

315The response always includes a `pagination_token`. When the returned page is shorter than `limit`, you've reached the end of the list.315When more files remain, the response includes a `pagination_token`. On the last page it is `null` or empty.

316 316 

317```pythonXAI317```pythonXAI

318import os318import os


399 399 

400### Paginating Through All Files400### Paginating Through All Files

401 401 

402The List endpoint returns at most `limit` files per call (capped at 100). To enumerate every file, keep calling the endpoint with the `pagination_token` from the previous response until the response returns fewer than `limit` items.402The List endpoint returns at most `limit` files per call (capped at 100). To enumerate every file, keep calling the endpoint with the `pagination_token` from the previous response until the response no longer includes one.

403 403 

404```pythonXAI404```pythonXAI

405import os405import os


407 407 

408client = Client(api_key=os.getenv("XAI_API_KEY"))408client = Client(api_key=os.getenv("XAI_API_KEY"))

409 409 

410# Walk every page until the API returns a short page.410# Walk every page until the API stops returning a pagination token.

411page_size = 100411page_size = 100

412token = None412token = None

413all_files = []413all_files = []


420 pagination_token=token,420 pagination_token=token,

421 )421 )

422 all_files.extend(response.data)422 all_files.extend(response.data)

423 if len(response.data) < page_size:

424 break

425 token = response.pagination_token423 token = response.pagination_token

424 if not token:

425 break

426 426 

427print(f"Total files: {len(all_files)}")427print(f"Total files: {len(all_files)}")

428```428```


441while True:441while True:

442 response = requests.get(url, headers=headers, params=params).json()442 response = requests.get(url, headers=headers, params=params).json()

443 all_files.extend(response.get("data", []))443 all_files.extend(response.get("data", []))

444 if len(response.get("data", [])) < page_size:444 if not response.get("pagination_token"):

445 break445 break

446 params["pagination_token"] = response["pagination_token"]446 params["pagination_token"] = response["pagination_token"]

447 447 


461 });461 });

462 const page = await response.json();462 const page = await response.json();

463 allFiles.push(...page.data);463 allFiles.push(...page.data);

464 if (page.data.length < pageSize) break;

465 token = page.pagination_token;464 token = page.pagination_token;

465 if (!token) break;

466}466}

467 467 

468console.log(\`Total files: \${allFiles.length}\`);468console.log(\`Total files: \${allFiles.length}\`);

grok-4-7.md +1 −1

Details

81| Input price | $2.00 / 1M tokens |81| Input price | $2.00 / 1M tokens |

82| Output price | $6.00 / 1M tokens |82| Output price | $6.00 / 1M tokens |

83| Reasoning | Low, medium, high (default), or xhigh |83| Reasoning | Low, medium, high (default), or xhigh |

84| APIs | [Responses API](/developers/rest-api-reference/inference/responses#create-new-response), [Chat Completions](/developers/rest-api-reference/inference/chat-completions#chat-completions) |84| APIs | [Responses API](/developers/rest-api-reference/inference/responses#create-new-response) |

85| Tools | [Function calling](/developers/tools/function-calling), [web search](/developers/tools/web-search), [X search](/developers/tools/x-search), [code execution](/developers/tools/code-execution) |85| Tools | [Function calling](/developers/tools/function-calling), [web search](/developers/tools/web-search), [X search](/developers/tools/x-search), [code execution](/developers/tools/code-execution) |

86 86 

87Rate limits and live pricing for your team are on the [model detail page](/developers/models/grok-4.7) and [Pricing](/developers/pricing).87Rate limits and live pricing for your team are on the [model detail page](/developers/models/grok-4.7) and [Pricing](/developers/pricing).

Details

61To retrieve a list of API keys from a team, you can run the following:61To retrieve a list of API keys from a team, you can run the following:

62 62 

63```bash63```bash

64curl https://management-api.x.ai/auth/teams/{teamId}/api-keys?pageSize=10&paginationToken= \\64curl "https://management-api.x.ai/auth/teams/{teamId}/api-keys?pageSize=10" \\

65 -H "Authorization: Bearer <Your Management API Key>"65 -H "Authorization: Bearer <Your Management API Key>"

66```66```

67 67 

68You can customize the query parameters such as `pageSize` and `paginationToken`.68Omit `paginationToken` on the first page; an empty value is rejected. To get the next page, pass the `paginationToken` from the previous response:

69 

70```bash

71curl "https://management-api.x.ai/auth/teams/{teamId}/api-keys?pageSize=10&paginationToken={paginationToken}" \\

72 -H "Authorization: Bearer <Your Management API Key>"

73```

74 

75Keep requesting pages until a response returns fewer than `pageSize` keys.

69 76 

70### Update an API key77### Update an API key

71 78 

Details

86 86 

87## Supported Languages87## Supported Languages

88 88 

89The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.89`grok-voice-transcribe-2.0` transcribes 38+ languages, including every language listed below. It detects the spoken language automatically and follows switches partway through a recording, so `language` is optional. When you know what's being spoken, set `language` to its code and the model leans toward that language when the audio is ambiguous.

90 90 

91| Language | Code | | Language | Code |91| Language | Code | | Language | Code |

92|----------|------|-|----------|------|92|----------|------|-|----------|------|

93| Arabic | `ar` | | Macedonian | `mk` |93| Arabic | `ar` | | Italian | `it` |

94| Czech | `cs` | | Malay | `ms` |94| Bosnian | `bs` | | Japanese | `ja` |

95| Danish | `da` | | Persian | `fa` |95| Bulgarian | `bg` | | Korean | `ko` |

96| Dutch | `nl` | | Polish | `pl` |96| Cantonese | `yue` | | Macedonian | `mk` |

97| English | `en` | | Portuguese | `pt` |97| Catalan | `ca` | | Malay | `ms` |

98| Filipino | `fil` | | Romanian | `ro` |98| Chinese (Mandarin) | `zh` | | Norwegian (Bokmål) | `nb` |

99| French | `fr` | | Russian | `ru` |99| Croatian | `hr` | | Persian | `fa` |

100| German | `de` | | Spanish | `es` |100| Czech | `cs` | | Polish | `pl` |

101| Hindi | `hi` | | Swedish | `sv` |101| Danish | `da` | | Portuguese | `pt` |

102| Indonesian | `id` | | Thai | `th` |102| Dutch | `nl` | | Romanian | `ro` |

103| Italian | `it` | | Turkish | `tr` |103| English | `en` | | Russian | `ru` |

104| Japanese | `ja` | | Vietnamese | `vi` |104| Filipino | `fil` | | Slovak | `sk` |

105| Korean | `ko` | | | |105| Finnish | `fi` | | Spanish | `es` |

106| French | `fr` | | Swedish | `sv` |

107| German | `de` | | Thai | `th` |

108| Greek | `el` | | Turkish | `tr` |

109| Hindi | `hi` | | Ukrainian | `uk` |

110| Hungarian | `hu` | | Urdu | `ur` |

111| Indonesian | `id` | | Vietnamese | `vi` |

112 

113Spanish, Portuguese, and Arabic also accept regional codes. `es` means Mexican Spanish (`es-MX`), `pt` means Brazilian Portuguese (`pt-BR`), and `ar` means Egyptian Arabic (`ar-EG`). Pass `es-ES` or `pt-PT` for the European variants, or `ar-AE` or `ar-SA` for Emirati or Saudi Arabic.

114 

115Text formatting (`format=true`) writes spoken numbers, currencies, and units in their written form in Arabic, Chinese (Mandarin), English, French, German, Japanese, Portuguese, Russian, Spanish, Swedish, and Vietnamese. Formatting follows the `language` code, so pass both. In other languages, `format=true` has no effect.

106 116 

107## Request Body117## Request Body

108 118 


115| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |125| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |

116| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |126| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |

117| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |127| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |

118| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |128| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Biases transcription toward that language and selects the formatting rules for `format=true`. See [Supported Languages](#supported-languages). |

119| `format` | boolean | `false` | | When `true`, enables Inverse Text Normalization — converts spoken numbers/currency to written form (e.g. "one hundred dollars" → "$100"). Requires `language`. |129| `format` | boolean | `false` | | When `true`, enables Inverse Text Normalization — converts spoken numbers/currency to written form (e.g. "one hundred dollars" → "$100"). Requires `language`. See [Supported Languages](#supported-languages) for coverage. |

120| `multichannel` | boolean | `false` | | When `true`, transcribes each audio channel independently. Results returned in the `channels` array. |130| `multichannel` | boolean | `false` | | When `true`, transcribes each audio channel independently. Results returned in the `channels` array. |

121| `channels` | integer | | | Number of audio channels (2–8). Only required for multichannel raw audio. Auto-detected for container formats. |131| `channels` | integer | | | Number of audio channels (2–8). Only required for multichannel raw audio. Auto-detected for container formats. |

122| `diarize` | boolean | `false` | | When `true`, enables speaker diarization. Each word in the response includes a `speaker` field (integer) identifying the detected speaker. |132| `diarize` | boolean | `false` | | When `true`, enables speaker diarization. Each word in the response includes a `speaker` field (integer) identifying the detected speaker. |


218| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |228| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |

219| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |229| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |

220| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |230| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |

221| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |231| `language` | string | | Language code (e.g. `en`, `fr`, `de`). Biases transcription toward that language and selects the formatting rules for `format=true`. See [Supported Languages](#supported-languages). |

232| `format` | boolean | `false` | When `true`, converts spoken numbers, currencies, and units to written form. Requires `language`. See [Supported Languages](#supported-languages) for coverage. |

222| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |233| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |

223| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |234| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |

224| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |235| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |


453 464 

454* **Use 16 kHz sample rate with PCM encoding** (`sample_rate=16000&encoding=pcm`) — this is the model's native rate and avoids resampling on the server465* **Use 16 kHz sample rate with PCM encoding** (`sample_rate=16000&encoding=pcm`) — this is the model's native rate and avoids resampling on the server

455* **Enable `interim_results`** for responsive UX — show transcription as the user speaks466* **Enable `interim_results`** for responsive UX — show transcription as the user speaks

456* **Use `language=en`** to enable text formatting — numbers and currencies are written in their standard form467* **Add `format=true` with `language`** (e.g. `language=en&format=true`) to write numbers and currencies in their standard form

457* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency468* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency

458* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s versus 48 KB/s for raw PCM at 24 kHz. See [Opus Streaming](#opus-streaming)469* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s versus 48 KB/s for raw PCM at 24 kHz. See [Opus Streaming](#opus-streaming)

459* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend470* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend

Details

132 132 

133## Speech to Text133## Speech to Text

134 134 

135Transcribe audio files in a single call or stream over WebSocket. The default is `grok-voice-transcribe-2.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.135Transcribe audio files in a single call or stream over WebSocket. The default is `grok-voice-transcribe-2.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 38+ languages.

136 136 

137```bash137```bash

138curl -X POST https://api.x.ai/v1/stt \138curl -X POST https://api.x.ai/v1/stt \

rate-limits.md +6 −5

Details

2 2 

3# Rate Limits3# Rate Limits

4 4 

5Every xAI API team has per-model rate limits on two dimensions: **requests per second (RPS)** and **tokens per minute (TPM)**. Your per-second limit is derived from your per-minute request budget (RPM / 60): you cannot spend a full minute's requests in a single second, which protects the API from sudden bursts. These limits scale with your team's **tier**, which is determined by cumulative spend on the API.5Every xAI API team has per-model rate limits on two dimensions: **requests per second (RPS)** and **tokens per minute (TPM)**. Your per-second limit is derived from your per-minute request budget (RPM / 48, with a minimum of 2): you cannot spend a full minute's requests in a single second, which protects the API from sudden bursts. These limits scale with your team's **tier**, which is determined by cumulative spend on the API.

6 6 

7You can view your team's current tier and per-model limits on the [Models](https://console.x.ai/team/default/models?utm_source=docs\&utm_medium=referral\&utm_campaign=developers-rate-limits\&utm_content=models) page in the xAI Console.7You can view your team's current tier and per-model limits on the [Models](https://console.x.ai/team/default/models?utm_source=docs\&utm_medium=referral\&utm_campaign=developers-rate-limits\&utm_content=models) page in the xAI Console.

8 8 


41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48| grok-imagine-video-1.5-lite | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

49| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |48| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

49| grok-imagine-video-1.5-lite | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

50 50 

51**Voice & Audio**51**Voice & Audio**

52 52 


86 messages=messages,86 messages=messages,

87 )87 )

88 except RateLimitError:88 except RateLimitError:

89 if attempt == max_retries - 1:

90 raise

89 wait = 2 ** attempt91 wait = 2 ** attempt

90 time.sleep(wait)92 time.sleep(wait)

91 raise RateLimitError("Max retries exceeded")

92```93```

93 94 

94```python customLanguage="pythonXAI"95```python customLanguage="pythonXAI"

Details

18 18 

19* `sample_rate` ("8000" | "16000" | "22050" | "24000" | "44100" | "48000") — Audio sample rate in Hz. \*\*Required when \`audio\_format\` is a raw format\*\* (\`pcm\`, \`mulaw\`, \`alaw\`). Ignored for container formats. Either \`sample\_rate\` or \`sample\_rate\_hertz\` may be used.19* `sample_rate` ("8000" | "16000" | "22050" | "24000" | "44100" | "48000") — Audio sample rate in Hz. \*\*Required when \`audio\_format\` is a raw format\*\* (\`pcm\`, \`mulaw\`, \`alaw\`). Ignored for container formats. Either \`sample\_rate\` or \`sample\_rate\_hertz\` may be used.

20 20 

21* `language` (string) — Language code for the audio (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set together with \`format=true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.21* `language` (string) — Language code for the audio (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). Biases transcription toward that language. When set together with \`format=true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form. Formatting is available for \`ar\`, \`de\`, \`en\`, \`es\`, \`fr\`, \`ja\`, \`pt\`, \`ru\`, \`sv\`, \`vi\`, and \`zh\`.

22 22 

23* `format` ("true" | "false") — When \`true\`, enables text formatting. Requires \`language\` to be set.23* `format` ("true" | "false") — When \`true\`, enables text formatting. Requires \`language\` to be set.

24 24 


154 154 

155* `endpointing` (integer, optional, default: 400) — Silence duration in milliseconds before the server fires a \`speech\_final=true\` event, indicating the speaker stopped talking. Range: 0–5000. Set to \`0\` for no delay (fire on any VAD silence boundary). Default: 400ms.155* `endpointing` (integer, optional, default: 400) — Silence duration in milliseconds before the server fires a \`speech\_final=true\` event, indicating the speaker stopped talking. Range: 0–5000. Set to \`0\` for no delay (fire on any VAD silence boundary). Default: 400ms.

156 156 

157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). Biases transcription toward that language and selects the formatting rules for \`format=true\`. See Supported Languages on the Speech to Text page.

158 

159* `format` (boolean, optional, default: false) — When \`true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form. Requires \`language\`; available for \`ar\`, \`de\`, \`en\`, \`es\`, \`fr\`, \`ja\`, \`pt\`, \`ru\`, \`sv\`, \`vi\`, and \`zh\`.

158 160 

159* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).161* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).

160 162 

Details

1016 1016 

1017* `sample_rate` ("8000" | "16000" | "22050" | "24000" | "44100" | "48000") — Audio sample rate in Hz. \*\*Required when \`audio\_format\` is a raw format\*\* (\`pcm\`, \`mulaw\`, \`alaw\`). Ignored for container formats. Either \`sample\_rate\` or \`sample\_rate\_hertz\` may be used.1017* `sample_rate` ("8000" | "16000" | "22050" | "24000" | "44100" | "48000") — Audio sample rate in Hz. \*\*Required when \`audio\_format\` is a raw format\*\* (\`pcm\`, \`mulaw\`, \`alaw\`). Ignored for container formats. Either \`sample\_rate\` or \`sample\_rate\_hertz\` may be used.

1018 1018 

1019* `language` (string) — Language code for the audio (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set together with \`format=true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.1019* `language` (string) — Language code for the audio (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). Biases transcription toward that language. When set together with \`format=true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form. Formatting is available for \`ar\`, \`de\`, \`en\`, \`es\`, \`fr\`, \`ja\`, \`pt\`, \`ru\`, \`sv\`, \`vi\`, and \`zh\`.

1020 1020 

1021* `format` ("true" | "false") — When \`true\`, enables text formatting. Requires \`language\` to be set.1021* `format` ("true" | "false") — When \`true\`, enables text formatting. Requires \`language\` to be set.

1022 1022 


1152 1152 

1153* `endpointing` (integer, optional, default: 400) — Silence duration in milliseconds before the server fires a \`speech\_final=true\` event, indicating the speaker stopped talking. Range: 0–5000. Set to \`0\` for no delay (fire on any VAD silence boundary). Default: 400ms.1153* `endpointing` (integer, optional, default: 400) — Silence duration in milliseconds before the server fires a \`speech\_final=true\` event, indicating the speaker stopped talking. Range: 0–5000. Set to \`0\` for no delay (fire on any VAD silence boundary). Default: 400ms.

1154 1154 

1155* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.1155* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). Biases transcription toward that language and selects the formatting rules for \`format=true\`. See Supported Languages on the Speech to Text page.

1156 

1157* `format` (boolean, optional, default: false) — When \`true\`, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form. Requires \`language\`; available for \`ar\`, \`de\`, \`en\`, \`es\`, \`fr\`, \`ja\`, \`pt\`, \`ru\`, \`sv\`, \`vi\`, and \`zh\`.

1156 1158 

1157* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).1159* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).

1158 1160