SpyBara
Go Premium

Documentation 2026-09-18 23:59 UTC to 2026-09-19 23:59 UTC

6 files changed +55 −13. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03
Details

82 82 

83| Model | Description |83| Model | Description |

84|-------|-------------|84|-------|-------------|

85| `grok-voice-transcribe-2.0` | Our best transcription model |85| `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |

86| `grok-voice-transcribe-1.0` | Original model. Default when `model` is omitted. |86| `grok-voice-transcribe-1.0` | Original model. Pin this slug to keep it. |

87 87 

88## Supported Languages88## Supported Languages

89 89 


113|-----------|------|---------|----------|-------------|113|-----------|------|---------|----------|-------------|

114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |

115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |

116| `model` | string | `grok-voice-transcribe-1.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |116| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |

117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |

118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |

119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |


220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |

221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |

222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |

223| `model` | string | `grok-voice-transcribe-1.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |223| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |

224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |

225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |

226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |

Details

132 132 

133## Speech to Text133## Speech to Text

134 134 

135Transcribe audio files in a single call or stream over WebSocket. Use `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`; the default is `grok-voice-transcribe-1.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.135Transcribe audio files in a single call or stream over WebSocket. Use `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`; the default is `grok-voice-transcribe-2.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.

136 136 

137```bash137```bash

138curl -X POST https://api.x.ai/v1/stt \138curl -X POST https://api.x.ai/v1/stt \

rate-limits.md +2 −2

Details

40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

43| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49**Voice & Audio**49**Voice & Audio**

50 50 

Details

16 16 

17 * `cached_prompt_text_token_price_long_context` (integer | null) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).17 * `cached_prompt_text_token_price_long_context` (integer | null) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).

18 18 

19 * `capabilities` (object)

20 

21 * `default_reasoning_effort` (string | null) — Model-wide default when effort is omitted; aliases may have their own defaults.

22 

23 * `reasoning_effort` (array\<string>, required) — The values the model accepts for \`reasoning\_effort\` or \`reasoning.effort\`.

24 

19 * `completion_text_token_price` (integer | null) — Price of the completion text token in USD cents per 100 million tokens.25 * `completion_text_token_price` (integer | null) — Price of the completion text token in USD cents per 100 million tokens.

20 26 

21 * `completion_text_token_price_long_context` (integer | null) — Price of the completion text token for long context requests (USD cents per 100 million tokens).27 * `completion_text_token_price_long_context` (integer | null) — Price of the completion text token for long context requests (USD cents per 100 million tokens).


87 "completion_text_token_price": 80000,93 "completion_text_token_price": 80000,

88 "prompt_text_token_price_long_context": 40000,94 "prompt_text_token_price_long_context": 40000,

89 "completion_text_token_price_long_context": 160000,95 "completion_text_token_price_long_context": 160000,

90 "long_context_threshold": 12800096 "long_context_threshold": 128000,

97 "capabilities": {

98 "reasoning_effort": [

99 "low",

100 "medium",

101 "high",

102 "xhigh"

103 ],

104 "default_reasoning_effort": "high"

105 }

91 },106 },

92 {107 {

93 "id": "grok-imagine-image",108 "id": "grok-imagine-image",


121 136 

122* `cached_prompt_text_token_price_long_context` (integer | null) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).137* `cached_prompt_text_token_price_long_context` (integer | null) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).

123 138 

139* `capabilities` (object)

140 

141 * `default_reasoning_effort` (string | null) — Model-wide default when effort is omitted; aliases may have their own defaults.

142 

143 * `reasoning_effort` (array\<string>, required) — The values the model accepts for \`reasoning\_effort\` or \`reasoning.effort\`.

144 

124* `completion_text_token_price` (integer | null) — Price of the completion text token in USD cents per 100 million tokens.145* `completion_text_token_price` (integer | null) — Price of the completion text token in USD cents per 100 million tokens.

125 146 

126* `completion_text_token_price_long_context` (integer | null) — Price of the completion text token for long context requests (USD cents per 100 million tokens).147* `completion_text_token_price_long_context` (integer | null) — Price of the completion text token for long context requests (USD cents per 100 million tokens).


192 * `cached_prompt_text_token_price_long_context` (integer, required) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).213 * `cached_prompt_text_token_price_long_context` (integer, required) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).

193 When 0, falls back to cached\_prompt\_text\_token\_price.214 When 0, falls back to cached\_prompt\_text\_token\_price.

194 215 

216 * `capabilities` (object)

217 

218 * `default_reasoning_effort` (string | null) — Model-wide default when effort is omitted; aliases may have their own defaults.

219 

220 * `reasoning_effort` (array\<string>, required) — The values the model accepts for \`reasoning\_effort\` or \`reasoning.effort\`.

221 

195 * `completion_text_token_price` (integer, required) — Price of the completion text token in USD cents per 100 million token.222 * `completion_text_token_price` (integer, required) — Price of the completion text token in USD cents per 100 million token.

196 223 

197 * `completion_text_token_price_long_context` (integer, required) — Price of the completion text token for long context requests (USD cents per 100 million tokens).224 * `completion_text_token_price_long_context` (integer, required) — Price of the completion text token for long context requests (USD cents per 100 million tokens).


279 "cached_prompt_text_token_price_long_context": 0,306 "cached_prompt_text_token_price_long_context": 0,

280 "completion_text_token_price_long_context": 160000,307 "completion_text_token_price_long_context": 160000,

281 "long_context_threshold": 128000,308 "long_context_threshold": 128000,

282 "aliases": []309 "aliases": [],

310 "capabilities": {

311 "reasoning_effort": [

312 "low",

313 "medium",

314 "high",

315 "xhigh"

316 ],

317 "default_reasoning_effort": "high"

318 }

283 }319 }

284 ]320 ]

285}321}


304* `cached_prompt_text_token_price_long_context` (integer, required) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).340* `cached_prompt_text_token_price_long_context` (integer, required) — Price of the cached prompt text token for long context requests (USD cents per 100 million tokens).

305 When 0, falls back to cached\_prompt\_text\_token\_price.341 When 0, falls back to cached\_prompt\_text\_token\_price.

306 342 

343* `capabilities` (object)

344 

345 * `default_reasoning_effort` (string | null) — Model-wide default when effort is omitted; aliases may have their own defaults.

346 

347 * `reasoning_effort` (array\<string>, required) — The values the model accepts for \`reasoning\_effort\` or \`reasoning.effort\`.

348 

307* `completion_text_token_price` (integer, required) — Price of the completion text token in USD cents per 100 million token.349* `completion_text_token_price` (integer, required) — Price of the completion text token in USD cents per 100 million token.

308 350 

309* `completion_text_token_price_long_context` (integer, required) — Price of the completion text token for long context requests (USD cents per 100 million tokens).351* `completion_text_token_price_long_context` (integer, required) — Price of the completion text token for long context requests (USD cents per 100 million tokens).

Details

140 140 

141WebSocket endpoint: `wss://api.x.ai/v1/stt`141WebSocket endpoint: `wss://api.x.ai/v1/stt`

142 142 

143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-1.0.143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-2.0.

144 144 

145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

146 146 


156 156 

157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

158 158 

159* `model` (string, optional, default: grok-voice-transcribe-1.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-1.0\`.159* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-2.0\`.

160 160 

161* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.161* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

162 162 

Details

1058 1058 

1059WebSocket endpoint: `wss://api.x.ai/v1/stt`1059WebSocket endpoint: `wss://api.x.ai/v1/stt`

1060 1060 

1061Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-1.0.1061Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-2.0.

1062 1062 

1063Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)1063Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

1064 1064 


1074 1074 

1075* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.1075* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

1076 1076 

1077* `model` (string, optional, default: grok-voice-transcribe-1.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-1.0\`.1077* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-2.0\`.

1078 1078 

1079* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.1079* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

1080 1080