86 86
87## Supported Languages87## Supported Languages
88 88
89The `language` parameter enables formatting for the following languages. The model transcribes speech in any of these languages regardless of the `language` parameter — setting it enables formatting of numbers, currencies, and units into their written form.89`grok-voice-transcribe-2.0` transcribes 38+ languages, including every language listed below. It detects the spoken language automatically and follows switches partway through a recording, so `language` is optional. When you know what's being spoken, set `language` to its code and the model leans toward that language when the audio is ambiguous.
90 90
91| Language | Code | | Language | Code |91| Language | Code | | Language | Code |
92|----------|------|-|----------|------|92|----------|------|-|----------|------|
93| Arabic | `ar` | | Macedonian | `mk` |93| Arabic | `ar` | | Italian | `it` |
94| Czech | `cs` | | Malay | `ms` |94| Bosnian | `bs` | | Japanese | `ja` |
95| Danish | `da` | | Persian | `fa` |95| Bulgarian | `bg` | | Korean | `ko` |
96| Dutch | `nl` | | Polish | `pl` |96| Cantonese | `yue` | | Macedonian | `mk` |
97| English | `en` | | Portuguese | `pt` |97| Catalan | `ca` | | Malay | `ms` |
98| Filipino | `fil` | | Romanian | `ro` |98| Chinese (Mandarin) | `zh` | | Norwegian (Bokmål) | `nb` |
99| French | `fr` | | Russian | `ru` |99| Croatian | `hr` | | Persian | `fa` |
100| German | `de` | | Spanish | `es` |100| Czech | `cs` | | Polish | `pl` |
101| Hindi | `hi` | | Swedish | `sv` |101| Danish | `da` | | Portuguese | `pt` |
102| Indonesian | `id` | | Thai | `th` |102| Dutch | `nl` | | Romanian | `ro` |
103| Italian | `it` | | Turkish | `tr` |103| English | `en` | | Russian | `ru` |
104| Japanese | `ja` | | Vietnamese | `vi` |104| Filipino | `fil` | | Slovak | `sk` |
105| Korean | `ko` | | | |105| Finnish | `fi` | | Spanish | `es` |
106| French | `fr` | | Swedish | `sv` |
107| German | `de` | | Thai | `th` |
108| Greek | `el` | | Turkish | `tr` |
109| Hindi | `hi` | | Ukrainian | `uk` |
110| Hungarian | `hu` | | Urdu | `ur` |
111| Indonesian | `id` | | Vietnamese | `vi` |
112
113Spanish, Portuguese, and Arabic also accept regional codes. `es` means Mexican Spanish (`es-MX`), `pt` means Brazilian Portuguese (`pt-BR`), and `ar` means Egyptian Arabic (`ar-EG`). Pass `es-ES` or `pt-PT` for the European variants, or `ar-AE` or `ar-SA` for Emirati or Saudi Arabic.
114
115Text formatting (`format=true`) writes spoken numbers, currencies, and units in their written form in Arabic, Chinese (Mandarin), English, French, German, Japanese, Portuguese, Russian, Spanish, Swedish, and Vietnamese. Formatting follows the `language` code, so pass both. In other languages, `format=true` has no effect.
106 116
107## Request Body117## Request Body
108 118
115| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |125| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |
116| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |126| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |
117| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |127| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |
118| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |128| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Biases transcription toward that language and selects the formatting rules for `format=true`. See [Supported Languages](#supported-languages). |
119| `format` | boolean | `false` | | When `true`, enables Inverse Text Normalization — converts spoken numbers/currency to written form (e.g. "one hundred dollars" → "$100"). Requires `language`. |129| `format` | boolean | `false` | | When `true`, enables Inverse Text Normalization — converts spoken numbers/currency to written form (e.g. "one hundred dollars" → "$100"). Requires `language`. See [Supported Languages](#supported-languages) for coverage. |
120| `multichannel` | boolean | `false` | | When `true`, transcribes each audio channel independently. Results returned in the `channels` array. |130| `multichannel` | boolean | `false` | | When `true`, transcribes each audio channel independently. Results returned in the `channels` array. |
121| `channels` | integer | | | Number of audio channels (2–8). Only required for multichannel raw audio. Auto-detected for container formats. |131| `channels` | integer | | | Number of audio channels (2–8). Only required for multichannel raw audio. Auto-detected for container formats. |
122| `diarize` | boolean | `false` | | When `true`, enables speaker diarization. Each word in the response includes a `speaker` field (integer) identifying the detected speaker. |132| `diarize` | boolean | `false` | | When `true`, enables speaker diarization. Each word in the response includes a `speaker` field (integer) identifying the detected speaker. |
218| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |228| `encoding` | string | `pcm` | Audio encoding: `pcm`, `mulaw`, `alaw`, or `opus`. See [Opus Streaming](#opus-streaming). |
219| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |229| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |
220| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |230| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |
221| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |231| `language` | string | | Language code (e.g. `en`, `fr`, `de`). Biases transcription toward that language and selects the formatting rules for `format=true`. See [Supported Languages](#supported-languages). |
232| `format` | boolean | `false` | When `true`, converts spoken numbers, currencies, and units to written form. Requires `language`. See [Supported Languages](#supported-languages) for coverage. |
222| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |233| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |
223| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |234| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |
224| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |235| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |
453 464
454* **Use 16 kHz sample rate with PCM encoding** (`sample_rate=16000&encoding=pcm`) — this is the model's native rate and avoids resampling on the server465* **Use 16 kHz sample rate with PCM encoding** (`sample_rate=16000&encoding=pcm`) — this is the model's native rate and avoids resampling on the server
455* **Enable `interim_results`** for responsive UX — show transcription as the user speaks466* **Enable `interim_results`** for responsive UX — show transcription as the user speaks
456* **Use `language=en`** to enable text formatting — numbers and currencies are written in their standard form467* **Add `format=true` with `language`** (e.g. `language=en&format=true`) to write numbers and currencies in their standard form
457* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency468* **Send 100 ms audio chunks** (3,200 bytes at 16 kHz PCM16) for a good balance of latency and efficiency
458* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s versus 48 KB/s for raw PCM at 24 kHz. See [Opus Streaming](#opus-streaming)469* **Use `encoding=opus` on bandwidth-constrained clients** — ~4 KB/s versus 48 KB/s for raw PCM at 24 kHz. See [Opus Streaming](#opus-streaming)
459* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend470* **Wait for `transcript.created`** before sending audio — the server needs to initialize its ASR backend