82 82
83| Model | Description |83| Model | Description |
84|-------|-------------|84|-------|-------------|
8585| `grok-voice-transcribe-2.0` | Our best transcription model || `grok-voice-transcribe-2.0` | Our best transcription model. Default when `model` is omitted. |
8686| `grok-voice-transcribe-1.0` | Original model. Default when `model` is omitted. || `grok-voice-transcribe-1.0` | Original model. Pin this slug to keep it. |
87 87
88## Supported Languages88## Supported Languages
89 89
113|-----------|------|---------|----------|-------------|113|-----------|------|---------|----------|-------------|
114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |114| `file` | file | | ✓† | Audio file to transcribe. Max **500 MB**. See [Supported Formats](#supported-audio-formats). Must be the last field in the multipart form. |
115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |115| `url` | string | | ✓† | URL of an audio file to download and transcribe (server-side). |
116116| `model` | string | `grok-voice-transcribe-1.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. || `model` | string | `grok-voice-transcribe-2.0` | | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |
117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |117| `audio_format` | string | | | Format hint for raw/headerless audio: `pcm`, `mulaw`, `alaw`. Container formats are auto-detected — do not set this field for MP3, WAV, etc. |
118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |118| `sample_rate` | integer | | | Sample rate in Hz. Only required for raw audio (`pcm`, `mulaw`, `alaw`). Supported: `8000`, `16000`, `22050`, `24000`, `44100`, `48000`. |
119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |119| `language` | string | | | Language code (e.g. `en`, `fr`, `de`). Used with `format=true` to enable text formatting. See [Supported Languages](#supported-languages). |
220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |220| `interim_results` | boolean | `false` | When `true`, emit partial transcripts `is_final=false` every ~500 ms. |
221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |221| `endpointing` | integer | `400` | Silence duration (ms) before utterance-final event. Range: 0–5000. `0` = fire on any VAD silence boundary. |
222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |222| `language` | string | | Language code for text formatting. See [Supported Languages](#supported-languages). |
223223| `model` | string | `grok-voice-transcribe-1.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. || `model` | string | `grok-voice-transcribe-2.0` | `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`. |
224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |224| `diarize` | boolean | | When `true`, enables speaker diarization. Words include a `speaker` field identifying the detected speaker. |
225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |225| `filler_words` | boolean | `false` | When `true`, filler words (e.g. `uh`, `um`, `er`) are included in the transcript. When `false` (default), filler words are automatically removed. |
226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |226| `multichannel` | boolean | `false` | Per-channel transcription. Requires `channels` ≥ 2. Not supported with `encoding=opus`. |