SpyBara
Go Premium

Documentation 2026-09-17 10:04 UTC to 2026-09-18 23:59 UTC

11 files changed +74 −33. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03

community.md +1 −1

Details

69 69 

70```javascriptAISDK70```javascriptAISDK

71import { xai } from '@ai-sdk/xai';71import { xai } from '@ai-sdk/xai';

72import { experimental_generateImage as generateImage } from 'ai';72import { generateImage } from 'ai';

73 73 

74const { image } = await generateImage({74const { image } = await generateImage({

75 model: xai.image('grok-imagine-image-2.0'),75 model: xai.image('grok-imagine-image-2.0'),

Details

11```bash11```bash

12curl -X POST https://api.x.ai/v1/stt \12curl -X POST https://api.x.ai/v1/stt \

13 -H "Authorization: Bearer $XAI_API_KEY" \13 -H "Authorization: Bearer $XAI_API_KEY" \

14 -F model=grok-voice-transcribe-2.0 \

14 -F format=true \15 -F format=true \

15 -F language=en \16 -F language=en \

16 -F "keyterm=Understand The Universe" \17 -F "keyterm=Understand The Universe" \


26 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},27 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},

27 files={"file": ("audio.mp3", open("audio.mp3", "rb"), "audio/mpeg")},28 files={"file": ("audio.mp3", open("audio.mp3", "rb"), "audio/mpeg")},

28 data=[29 data=[

30 ("model", "grok-voice-transcribe-2.0"),

29 ("format", "true"),31 ("format", "true"),

30 ("language", "en"),32 ("language", "en"),

31 ("keyterm", "Understand The Universe"),33 ("keyterm", "Understand The Universe"),


44import fs from "fs";46import fs from "fs";

45 47 

46const formData = new FormData();48const formData = new FormData();

49formData.append("model", "grok-voice-transcribe-2.0");

47formData.append("format", "true");50formData.append("format", "true");

48formData.append("language", "en");51formData.append("language", "en");

49formData.append("keyterm", "Understand The Universe");52formData.append("keyterm", "Understand The Universe");


73 76 

74[Live Voice Demos](https://x.ai/api/voice)77[Live Voice Demos](https://x.ai/api/voice)

75 78 

79## Model Selection

80 

81Pass `model` on the REST form or as a WebSocket query parameter.

82 

83| Model | Description |

84|-------|-------------|

85| `grok-voice-transcribe-2.0` | Our best transcription model |

86| `grok-voice-transcribe-1.0` | Original model. Default when `model` is omitted. |

87 

76## Supported Languages88## Supported Languages

77 89 

78The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.90The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.


101|-----------|------|---------|----------|-------------|113|-----------|------|---------|----------|-------------|

102| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |

103| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |

116| `model` | string | `grok-voice-transcribe-1.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |

104| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |

105| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |

106| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |


121```bash134```bash

122curl -X POST https://api.x.ai/v1/stt \135curl -X POST https://api.x.ai/v1/stt \

123 -H "Authorization: Bearer $XAI_API_KEY" \136 -H "Authorization: Bearer $XAI_API_KEY" \

137 -F model=grok-voice-transcribe-2.0 \

124 -F format=true \138 -F format=true \

125 -F language=en \139 -F language=en \

126 -F "keyterm=Understand The Universe" \140 -F "keyterm=Understand The Universe" \


206| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |

207| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |

208| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |

223| `model` | string | `grok-voice-transcribe-1.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |

209| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |

210| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |

211| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |


267**Example URL:**282**Example URL:**

268 283 

269```284```

270wss://api.x.ai/v1/stt?encoding=opus&interim_results=true285wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&encoding=opus&interim_results=true

271```286```

272 287 

273**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.288**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.


287**Example URL:**302**Example URL:**

288 303 

289```304```

290wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&multichannel=true&channels=2&interim_results=true305wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&multichannel=true&channels=2&interim_results=true

291```306```

292 307 

293**Typical use case:** Call center recordings with agent on channel 0 and customer on channel 1, enabling per-speaker transcription without requiring speaker diarization.308**Typical use case:** Call center recordings with agent on channel 0 and customer on channel 1, enabling per-speaker transcription without requiring speaker diarization.


315When Smart Turn is enabled, the model has full control over when `speech_final` fires. To prevent sessions from hanging during extended silence (e.g. the user walks away), set `smart_turn_timeout` to a maximum silence duration in milliseconds (1–5000). If the model keeps predicting "not done" for longer than this duration, `speech_final` fires anyway as a safety net.330When Smart Turn is enabled, the model has full control over when `speech_final` fires. To prevent sessions from hanging during extended silence (e.g. the user walks away), set `smart_turn_timeout` to a maximum silence duration in milliseconds (1–5000). If the model keeps predicting "not done" for longer than this duration, `speech_final` fires anyway as a safety net.

316 331 

317```332```

318wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&smart_turn=0.7&smart_turn_timeout=3000333wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&smart_turn=0.7&smart_turn_timeout=3000

319```334```

320 335 

321Without `smart_turn_timeout`, the model has unlimited control — `speech_final` only fires when confidence exceeds the threshold.336Without `smart_turn_timeout`, the model has unlimited control — `speech_final` only fires when confidence exceeds the threshold.


347import websockets362import websockets

348 363 

349API_KEY = os.environ["XAI_API_KEY"]364API_KEY = os.environ["XAI_API_KEY"]

350WS_URL = "wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe"365WS_URL = "wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe"

351 366 

352async def transcribe_stream(audio_file: str):367async def transcribe_stream(audio_file: str):

353 headers = {"Authorization": f"Bearer {API_KEY}"}368 headers = {"Authorization": f"Bearer {API_KEY}"}


389import WebSocket from "ws";404import WebSocket from "ws";

390 405 

391const apiKey = process.env.XAI_API_KEY;406const apiKey = process.env.XAI_API_KEY;

392const url = "wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe";407const url = "wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe";

393 408 

394const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${apiKey}` } });409const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${apiKey}` } });

395 410 

Details

132 132 

133## Speech to Text133## Speech to Text

134 134 

135Transcribe audio files in a single call or stream over WebSocket. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.135Transcribe audio files in a single call or stream over WebSocket. Use `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`; the default is `grok-voice-transcribe-1.0`. 12 audio formats, word-level timestamps, multichannel, speaker diarization, Smart Turn end-of-turn detection, and 25 languages.

136 136 

137```bash137```bash

138curl -X POST https://api.x.ai/v1/stt \138curl -X POST https://api.x.ai/v1/stt \

139 -H "Authorization: Bearer $XAI_API_KEY" \139 -H "Authorization: Bearer $XAI_API_KEY" \

140 -F model=grok-voice-transcribe-2.0 \

140 -F file=@recording.mp3141 -F file=@recording.mp3

141```142```

142 143 


148 "https://api.x.ai/v1/stt",149 "https://api.x.ai/v1/stt",

149 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},150 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},

150 files={"file": ("recording.mp3", open("recording.mp3", "rb"), "audio/mpeg")},151 files={"file": ("recording.mp3", open("recording.mp3", "rb"), "audio/mpeg")},

152 data={"model": "grok-voice-transcribe-2.0"},

151)153)

152 154 

153print(response.json()["text"])155print(response.json()["text"])


157import fs from "fs";159import fs from "fs";

158 160 

159const formData = new FormData();161const formData = new FormData();

162formData.append("model", "grok-voice-transcribe-2.0");

160formData.append("file", new Blob([fs.readFileSync("recording.mp3")]), "recording.mp3");163formData.append("file", new Blob([fs.readFileSync("recording.mp3")]), "recording.mp3");

161 164 

162const response = await fetch("https://api.x.ai/v1/stt", {165const response = await fetch("https://api.x.ai/v1/stt", {

Details

23 23 

24client = Client(24client = Client(

25 api_key=os.getenv('XAI_API_KEY'),25 api_key=os.getenv('XAI_API_KEY'),

26 timeout=3600, # Override default timeout with longer timeout for reasoning models26 timeout=3600,

27)27)

28 28 

29chat = client.chat.create(model="grok-4.6")29chat = client.chat.create(model="grok-4.6")


35)35)

36 36 

37for response, chunk in chat.stream():37for response, chunk in chat.stream():

38 print(chunk.content, end="", flush=True) # Each chunk's content38 print(chunk.content, end="", flush=True)

39 print(response.content, end="", flush=True) # The response object auto-accumulates the chunks

40 39 

41print(response.content) # The full response40print()

41print(response.content)

42```42```

43 43 

44```pythonOpenAISDK44```pythonOpenAISDK


49client = OpenAI(49client = OpenAI(

50 api_key=os.getenv("XAI_API_KEY"),50 api_key=os.getenv("XAI_API_KEY"),

51 base_url="https://api.x.ai/v1",51 base_url="https://api.x.ai/v1",

52 timeout=httpx.Timeout(3600.0) # Timeout after 3600s for reasoning models52 timeout=httpx.Timeout(3600.0)

53)53)

54 54 

55stream = client.chat.completions.create(55stream = client.chat.completions.create(


58 {"role": "system", "content": "You are Grok, a helpful and useful AI built by xAI."},58 {"role": "system", "content": "You are Grok, a helpful and useful AI built by xAI."},

59 {"role": "user", "content": "Explain how neural networks learn in two sentences."},59 {"role": "user", "content": "Explain how neural networks learn in two sentences."},

60 ],60 ],

61 stream=True # Set streaming here61 stream=True

62)62)

63 63 

64for chunk in stream:64for chunk in stream:

65 if chunk.choices[0].delta.content:

65 print(chunk.choices[0].delta.content, end="", flush=True)66 print(chunk.choices[0].delta.content, end="", flush=True)

66```67```

67 68 

68```javascriptOpenAISDK69```javascriptOpenAISDK

69import OpenAI from "openai";70import OpenAI from "openai";

70const openai = new OpenAI({71const openai = new OpenAI({

71 apiKey: "<api key>",72 apiKey: process.env.XAI_API_KEY,

72 baseURL: "https://api.x.ai/v1",73 baseURL: "https://api.x.ai/v1",

73 timeout: 360000, // Timeout after 3600s for reasoning models74 timeout: 360000,

74});75});

75 76 

76const stream = await openai.chat.completions.create({77const stream = await openai.chat.completions.create({


86});87});

87 88 

88for await (const chunk of stream) {89for await (const chunk of stream) {

89 console.log(chunk.choices[0].delta.content);90 const content = chunk.choices[0].delta.content;

91 if (content) {

92 process.stdout.write(content);

93 }

90}94}

91```95```

92 96 


130You'll get the event streams like these:134You'll get the event streams like these:

131 135 

132```json136```json

137data: {

138 "id":"<completion_id>","object":"chat.completion.chunk","created":<creation_time>,

139 "model":"grok-4.6",

140 "choices":[{"index":0,"delta":{"reasoning_content":"The","role":"assistant"}}],

141 "usage":{"prompt_tokens":41,"completion_tokens":1,"total_tokens":42,

142 "prompt_tokens_details":{"text_tokens":41,"audio_tokens":0,"image_tokens":0,"cached_tokens":0}},

143 "system_fingerprint":"fp_xxxxxxxxxx"

144}

145 

133data: {146data: {

134 "id":"<completion_id>","object":"chat.completion.chunk","created":<creation_time>,147 "id":"<completion_id>","object":"chat.completion.chunk","created":<creation_time>,

135 "model":"grok-4.6",148 "model":"grok-4.6",

rate-limits.md +7 −5

Details

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |43| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49**Voice & Audio**49**Voice & Audio**

50 50 


92```python customLanguage="pythonXAI"92```python customLanguage="pythonXAI"

93import os93import os

94import time94import time

95import grpc

95from xai_sdk import Client96from xai_sdk import Client

96from xai_sdk.chat import user97from xai_sdk.chat import user

97from xai_sdk.exceptions import RateLimitError

98 98 

99client = Client(api_key=os.getenv("XAI_API_KEY"))99client = Client(api_key=os.getenv("XAI_API_KEY"))

100 100 


104 for attempt in range(max_retries):104 for attempt in range(max_retries):

105 try:105 try:

106 return chat.sample()106 return chat.sample()

107 except RateLimitError:107 except grpc.RpcError as e:

108 if e.code() != grpc.StatusCode.RESOURCE_EXHAUSTED:

109 raise

108 wait = 2 ** attempt110 wait = 2 ** attempt

109 time.sleep(wait)111 time.sleep(wait)

110 raise RateLimitError("Max retries exceeded")112 raise RuntimeError("Max retries exceeded")

111```113```

112 114 

113## Increasing your limits115## Increasing your limits

Details

42 42 

43* `response_format` (object | object | object)43* `response_format` (object | object | object)

44 44 

45* `safety_identifier` (string | null) — Supplied by the API client to identify the end user behind this request. A stable string that uniquely identifies each of your users; hash your internal user id or username rather than sending an email or name. Stored with the request metadata so a usage-policy violation can be attributed to that user rather than to the API key.

46 

45* `search_parameters` (object)47* `search_parameters` (object)

46 48 

47 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See49 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

Details

16 16 

17* `prompt` (string) — Prompt for image generation.17* `prompt` (string) — Prompt for image generation.

18 18 

19* `resolution` ("1k" | "2k")19* `resolution` ("1k" | "2k" | "1.5k")

20 20 

21* `response_format` (string | null) — Response format to return the image in. Can be url or b64\_json. If b64\_json is specified, the image will be returned as a base64-encoded string instead of a url to the generated image file.21* `response_format` (string | null) — Response format to return the image in. Can be url or b64\_json. If b64\_json is specified, the image will be returned as a base64-encoded string instead of a url to the generated image file.

22 22 


195 195 

196* `prompt` (string, required) — Prompt for image editing.196* `prompt` (string, required) — Prompt for image editing.

197 197 

198* `resolution` ("1k" | "2k")198* `resolution` ("1k" | "2k" | "1.5k")

199 199 

200* `response_format` (string | null) — Response format to return the image in. Can be \`url\` or \`b64\_json\`. If \`b64\_json\` is specified, the image will be returned as a base64-encoded string instead of a url to the generated image file.200* `response_format` (string | null) — Response format to return the image in. Can be \`url\` or \`b64\_json\`. If \`b64\_json\` is specified, the image will be returned as a base64-encoded string instead of a url to the generated image file.

201 201 

Details

46 Medium is the default quality a request serves at when it leaves46 Medium is the default quality a request serves at when it leaves

47 \`quality\` unset.47 \`quality\` unset.

48 48 

49 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the49 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\`, \`"1.5k"\`, or \`"2k"\`.

50 default when a request leaves \`resolution\` unset.50 1k is the default when a request leaves \`resolution\` unset.

51 51 

52 * `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.52 * `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.

53 53 


151 Medium is the default quality a request serves at when it leaves151 Medium is the default quality a request serves at when it leaves

152 \`quality\` unset.152 \`quality\` unset.

153 153 

154 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the154 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\`, \`"1.5k"\`, or \`"2k"\`.

155 default when a request leaves \`resolution\` unset.155 1k is the default when a request leaves \`resolution\` unset.

156 156 

157* `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.157* `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.

158 158 


406 Medium is the default quality a request serves at when it leaves406 Medium is the default quality a request serves at when it leaves

407 \`quality\` unset.407 \`quality\` unset.

408 408 

409 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the409 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\`, \`"1.5k"\`, or \`"2k"\`.

410 default when a request leaves \`resolution\` unset.410 1k is the default when a request leaves \`resolution\` unset.

411 411 

412 * `version` (string, required) — Version of the model.412 * `version` (string, required) — Version of the model.

413 413 


477 Medium is the default quality a request serves at when it leaves477 Medium is the default quality a request serves at when it leaves

478 \`quality\` unset.478 \`quality\` unset.

479 479 

480 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the480 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\`, \`"1.5k"\`, or \`"2k"\`.

481 default when a request leaves \`resolution\` unset.481 1k is the default when a request leaves \`resolution\` unset.

482 482 

483* `version` (string, required) — Version of the model.483* `version` (string, required) — Version of the model.

484 484 

Details

56 56 

57* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.57* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.

58 58 

59* `safety_identifier` (string | null) — Supplied by the API client to identify the end user behind this request. A stable string that uniquely identifies each of your users; hash your internal user id or username rather than sending an email or name. Stored with the request metadata so a usage-policy violation can be attributed to that user rather than to the API key.

60 

59* `search_parameters` (object)61* `search_parameters` (object)

60 62 

61 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See63 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

Details

140 140 

141WebSocket endpoint: `wss://api.x.ai/v1/stt`141WebSocket endpoint: `wss://api.x.ai/v1/stt`

142 142 

143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time.143Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-1.0.

144 144 

145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)145Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

146 146 


156 156 

157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.157* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

158 158 

159* `model` (string, optional, default: grok-voice-transcribe-1.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-1.0\`.

160 

159* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.161* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

160 162 

161* `channels` (integer, optional, default: 1) — Number of interleaved audio channels. Required when \`multichannel=true\`. Min: 2, Max: 8.163* `channels` (integer, optional, default: 1) — Number of interleaved audio channels. Required when \`multichannel=true\`. Min: 2, Max: 8.

Details

1058 1058 

1059WebSocket endpoint: `wss://api.x.ai/v1/stt`1059WebSocket endpoint: `wss://api.x.ai/v1/stt`

1060 1060 

1061Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time.1061Real-time streaming speech-to-text via WebSocket. Stream raw audio as binary frames and receive JSON transcript events as the audio is processed. Configuration is done via query parameters at connection time. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-1.0.

1062 1062 

1063Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)1063Full schemas and examples: [`/stt-streaming.ws.json`](/stt-streaming.ws.json)

1064 1064 


1074 1074 

1075* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.1075* `language` (string, optional, default: ) — Language code (e.g. \`en\`, \`fr\`, \`de\`, \`ja\`). When set, enables Inverse Text Normalization — spoken-form numbers, currencies, and units are converted to their written form.

1076 1076 

1077* `model` (string, optional, default: grok-voice-transcribe-1.0) — \`grok-voice-transcribe-1.0\` or \`grok-voice-transcribe-2.0\`. Defaults to \`grok-voice-transcribe-1.0\`.

1078 

1077* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.1079* `multichannel` (boolean, optional, default: false) — When \`true\`, enables per-channel transcription for interleaved multichannel audio. Requires \`channels\` to be set to ≥ 2. Not supported with \`encoding=opus\`.

1078 1080 

1079* `channels` (integer, optional, default: 1) — Number of interleaved audio channels. Required when \`multichannel=true\`. Min: 2, Max: 8.1081* `channels` (integer, optional, default: 1) — Number of interleaved audio channels. Required when \`multichannel=true\`. Min: 2, Max: 8.