83| Model | Description |83| Model | Description |
84|-------|-------------|84|-------|-------------|
85| `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |85| `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |
86| `grok-voice-transcribe-1.0` | Original model. Deprecated; end of life October 2, 2026. Requests are routed to `grok-voice-transcribe-2.0` at the same price, with higher accuracy. |
87 86
88## Supported Languages87## Supported Languages
89 88
113|-----------|------|---------|----------|-------------|112|-----------|------|---------|----------|-------------|
114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |113| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |
115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |114| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |
116115| `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default) or `grok-voice-transcribe-1.0` (deprecated; routed to 2.0). || `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-2.0` (default). |
117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |116| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |
118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |117| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |
119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |118| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |
220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |219| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |
221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |220| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |
222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |221| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |
223222| `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default) or `grok-voice-transcribe-1.0` (deprecated; routed to 2.0). || `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-2.0` (default). |
224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |223| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |
225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |224| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |
226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |225| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |