11```bash11```bash
12curl -X POST https://api.x.ai/v1/stt \12curl -X POST https://api.x.ai/v1/stt \
13 -H "Authorization: Bearer $XAI_API_KEY" \13 -H "Authorization: Bearer $XAI_API_KEY" \
14 -F model=grok-voice-transcribe-2.0 \
14 -F format=true \15 -F format=true \
15 -F language=en \16 -F language=en \
16 -F "keyterm=Understand The Universe" \17 -F "keyterm=Understand The Universe" \
26 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},27 headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
27 files={"file": ("audio.mp3", open("audio.mp3", "rb"), "audio/mpeg")},28 files={"file": ("audio.mp3", open("audio.mp3", "rb"), "audio/mpeg")},
28 data=[29 data=[
30 ("model", "grok-voice-transcribe-2.0"),
29 ("format", "true"),31 ("format", "true"),
30 ("language", "en"),32 ("language", "en"),
31 ("keyterm", "Understand The Universe"),33 ("keyterm", "Understand The Universe"),
44import fs from "fs";46import fs from "fs";
45 47
46const formData = new FormData();48const formData = new FormData();
49formData.append("model", "grok-voice-transcribe-2.0");
47formData.append("format", "true");50formData.append("format", "true");
48formData.append("language", "en");51formData.append("language", "en");
49formData.append("keyterm", "Understand The Universe");52formData.append("keyterm", "Understand The Universe");
73 76
74[Live Voice Demos](https://x.ai/api/voice)77[Live Voice Demos](https://x.ai/api/voice)
75 78
79## Model Selection
80
81Pass `model` on the REST form or as a WebSocket query parameter.
82
83| Model | Description |
84|-------|-------------|
85| `grok-voice-transcribe-2.0` | Our best transcription model |
86| `grok-voice-transcribe-1.0` | Original model. Default when `model` is omitted. |
87
76## Supported Languages88## Supported Languages
77 89
78The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.90The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.
101|-----------|------|---------|----------|-------------|113|-----------|------|---------|----------|-------------|
102| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |
103| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |
116| `model` | string | `grok-voice-transcribe-1.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |
104| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |
105| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |
106| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |
121```bash134```bash
122curl -X POST https://api.x.ai/v1/stt \135curl -X POST https://api.x.ai/v1/stt \
123 -H "Authorization: Bearer $XAI_API_KEY" \136 -H "Authorization: Bearer $XAI_API_KEY" \
137 -F model=grok-voice-transcribe-2.0 \
124 -F format=true \138 -F format=true \
125 -F language=en \139 -F language=en \
126 -F "keyterm=Understand The Universe" \140 -F "keyterm=Understand The Universe" \
206| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |
207| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |
208| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |
223| `model` | string | `grok-voice-transcribe-1.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |
209| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |
210| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |
211| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |
267**Example URL:**282**Example URL:**
268 283
269```284```
270wss://api.x.ai/v1/stt?encoding=opus&interim_results=true285wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&encoding=opus&interim_results=true
271```286```
272 287
273**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.288**Typical use case:** Dictation and live transcription from mobile or other bandwidth-constrained clients, where streaming raw PCM is wasteful. Most platform audio APIs and WebRTC stacks produce Opus packets natively.
287**Example URL:**302**Example URL:**
288 303
289```304```
290wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&multichannel=true&channels=2&interim_results=true305wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&multichannel=true&channels=2&interim_results=true
291```306```
292 307
293**Typical use case:** Call center recordings with agent on channel 0 and customer on channel 1, enabling per-speaker transcription without requiring speaker diarization.308**Typical use case:** Call center recordings with agent on channel 0 and customer on channel 1, enabling per-speaker transcription without requiring speaker diarization.
315When Smart Turn is enabled, the model has full control over when `speech_final` fires. To prevent sessions from hanging during extended silence (e.g. the user walks away), set `smart_turn_timeout` to a maximum silence duration in milliseconds (1–5000). If the model keeps predicting "not done" for longer than this duration, `speech_final` fires anyway as a safety net.330When Smart Turn is enabled, the model has full control over when `speech_final` fires. To prevent sessions from hanging during extended silence (e.g. the user walks away), set `smart_turn_timeout` to a maximum silence duration in milliseconds (1–5000). If the model keeps predicting "not done" for longer than this duration, `speech_final` fires anyway as a safety net.
316 331
317```332```
318wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&smart_turn=0.7&smart_turn_timeout=3000333wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&smart_turn=0.7&smart_turn_timeout=3000
319```334```
320 335
321Without `smart_turn_timeout`, the model has unlimited control — `speech_final` only fires when confidence exceeds the threshold.336Without `smart_turn_timeout`, the model has unlimited control — `speech_final` only fires when confidence exceeds the threshold.
347import websockets362import websockets
348 363
349API_KEY = os.environ["XAI_API_KEY"]364API_KEY = os.environ["XAI_API_KEY"]
350WS_URL = "wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe"365WS_URL = "wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe"
351 366
352async def transcribe_stream(audio_file: str):367async def transcribe_stream(audio_file: str):
353 headers = {"Authorization": f"Bearer {API_KEY}"}368 headers = {"Authorization": f"Bearer {API_KEY}"}
389import WebSocket from "ws";404import WebSocket from "ws";
390 405
391const apiKey = process.env.XAI_API_KEY;406const apiKey = process.env.XAI_API_KEY;
392const url = "wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe";407const url = "wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en&keyterm=Understand+The+Universe";
393 408
394const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${apiKey}` } });409const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${apiKey}` } });
395 410