SpyBara
Go Premium

resources/audio/subresources/translations/index.md 2026-07-10 23:02 UTC to 2026-07-12 06:58 UTC

1 added, 1 removed.

2026
Fri 31 21:03 Thu 30 23:58 Wed 29 15:02 Sat 25 05:59 Thu 23 18:00 Wed 22 20:02 Mon 20 20:00 Fri 17 17:00 Thu 16 20:57 Wed 15 02:58 Tue 14 06:58 Mon 13 15:59 Sun 12 06:58 Fri 10 23:02 Thu 9 20:58 Tue 7 08:02

Translations

Create translation

post /audio/translations

Translates audio into English.

Returns

  • Translation object { text }

    • text: string
  • TranslationVerbose object { duration, language, text, segments }

    • duration: number

      The duration of the input audio.

    • language: string

      The language of the output translation (always english).

    • text: string

      The translated text.

    • segments: optional array of TranscriptionSegment

      Segments of the translated text and their corresponding details.

      • id: number

        Unique identifier of the segment.

      • avg_logprob: number

        Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

      • compression_ratio: number

        Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

      • end: number

        End time of the segment in seconds.

      • no_speech_prob: number

        Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

      • seek: number

        Seek offset of the segment.

      • start: number

        Start time of the segment in seconds.

      • temperature: number

        Temperature parameter used for generating the segment.

      • text: string

        Text content of the segment.

      • tokens: array of number

        Array of token IDs for the text content.

Example

curl https://api.openai.com/v1/audio/translations \
    -H 'Content-Type: multipart/form-data' \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -F 'file=@/path/to/file' \
    -F model=whisper-1

Response

{
  "text": "text"
}

Example

curl https://api.openai.com/v1/audio/translations \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: multipart/form-data" \
  -F file="@/path/to/file/german.m4a" \
  -F model="whisper-1"

Response

{
  "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}

Domain Types

Translation

  • Translation object { text }

    • text: string

Translation Verbose

  • TranslationVerbose object { duration, language, text, segments }

    • duration: number

      The duration of the input audio.

    • language: string

      The language of the output translation (always english).

    • text: string

      The translated text.

    • segments: optional array of TranscriptionSegment

      Segments of the translated text and their corresponding details.

      • id: number

        Unique identifier of the segment.

      • avg_logprob: number

        Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

      • compression_ratio: number

        Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

      • end: number

        End time of the segment in seconds.

      • no_speech_prob: number

        Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

      • seek: number

        Seek offset of the segment.

      • start: number

        Start time of the segment in seconds.

      • temperature: number

        Temperature parameter used for generating the segment.

      • text: string

        Text content of the segment.

      • tokens: array of number

        Array of token IDs for the text content.

Translation Create Response

  • TranslationCreateResponse = Translation or TranslationVerbose

    • Translation object { text }

      • text: string
    • TranslationVerbose object { duration, language, text, segments }

      • duration: number

        The duration of the input audio.

      • language: string

        The language of the output translation (always english).

      • text: string

        The translated text.

      • segments: optional array of TranscriptionSegment

        Segments of the translated text and their corresponding details.

        • id: number

          Unique identifier of the segment.

        • avg_logprob: number

          Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

        • compression_ratio: number

          Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

        • end: number

          End time of the segment in seconds.

        • no_speech_prob: number

          Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

        • seek: number

          Seek offset of the segment.

        • start: number

          Start time of the segment in seconds.

        • temperature: number

          Temperature parameter used for generating the segment.

        • text: string

          Text content of the segment.

        • tokens: array of number

          Array of token IDs for the text content.