SpyBara
Go Premium

Documentation 2026-10-02 23:57 UTC to 2026-10-03 23:58 UTC

5 files changed +10 −11. View all changes and history on the product overview
2026
Sun 11 03:59 Sat 10 23:59 Fri 9 23:59 Thu 8 23:58 Wed 7 23:59 Tue 6 22:57 Mon 5 22:59 Sun 4 23:58 Sat 3 23:58 Fri 2 23:57 Thu 1 23:57
Details

83| Model | Description |83| Model | Description |

84|-------|-------------|84|-------|-------------|

85| `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |85| `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |

86| `grok-voice-transcribe-1.0` | Original model. Deprecated; end of life October 2, 2026. Requests are routed to `grok-voice-transcribe-2.0` at the same price, with higher accuracy. |

87 86 

88## Supported Languages87## Supported Languages

89 88 


113|-----------|------|---------|----------|-------------|112|-----------|------|---------|----------|-------------|

114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |113| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |

115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |114| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |

116| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default) or `grok-voice-transcribe-1.0` (deprecated; routed to 2.0). |115| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |

117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |116| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |

118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |117| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |

119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |118| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |


220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |219| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |

221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |220| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |

222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |221| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |

223| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default) or `grok-voice-transcribe-1.0` (deprecated; routed to 2.0). |222| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |

224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |223| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |

225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |224| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |

226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |225| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |

Details

132 132 

133## Speech to Text133## Speech to Text

134 134 

135Transcribe audio files in a single call or stream over WebSocket. The default is `grok-voice-transcribe-2.0`. `grok-voice-transcribe-1.0` is deprecated and reaches end of life on October 2, 2026; requests to that slug are routed to 2.0. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.135Transcribe audio files in a single call or stream over WebSocket. The default is `grok-voice-transcribe-2.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.

136 136 

137```bash137```bash

138curl -X POST https://api.x.ai/v1/stt \138curl -X POST https://api.x.ai/v1/stt \

rate-limits.md +3 −3

Details

41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |42| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |43| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video-1.5-lite | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

48| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

49| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |48| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

49| grok-imagine-video-1.5-lite | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

50 50 

51**Voice & Audio**51**Voice & Audio**

52 52 

Details

140 140 

141WebSocket endpoint: `wss://api.x.ai/v1/stt`141WebSocket endpoint: `wss://api.x.ai/v1/stt`

142 142 

143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. The default is grok-voice-transcribe-2.0. grok-voice-transcribe-1.0 is deprecated and reaches end of life on October 2, 2026; requests to that slug are routed to grok-voice-transcribe-2.0.143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. The default is grok-voice-transcribe-2.0.

144 144 

145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

146 146 


156 156 

157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

158 158 

159* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default) or \`grok-voice-transcribe-1.0\` (deprecated; routed to 2.0 as of October 2, 2026).159* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).

160 160 

161* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.161* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

162 162 

Details

1138 1138 

1139WebSocket endpoint: `wss://api.x.ai/v1/stt`1139WebSocket endpoint: `wss://api.x.ai/v1/stt`

1140 1140 

1141Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. The default is grok-voice-transcribe-2.0. grok-voice-transcribe-1.0 is deprecated and reaches end of life on October 2, 2026; requests to that slug are routed to grok-voice-transcribe-2.0.1141Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. The default is grok-voice-transcribe-2.0.

1142 1142 

1143Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)1143Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

1144 1144 


1154 1154 

1155* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.1155* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

1156 1156 

1157* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default) or \`grok-voice-transcribe-1.0\` (deprecated; routed to 2.0 as of October 2, 2026).1157* `model` (string, optional, default: grok-voice-transcribe-2.0) — \`grok-voice-transcribe-2.0\` (default).

1158 1158 

1159* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.1159* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

1160 1160