SpyBara
Go Premium

resources/audio/subresources/translations/index.md 2026-05-02 05:57 UTC to 2026-05-05 23:00 UTC

241 added, 0 removed.

2026
Wed 27 06:42 Fri 22 06:33 Wed 20 06:35 Tue 19 06:34 Mon 18 22:01 Mon 11 18:00 Thu 7 21:57 Tue 5 23:00 Sat 2 05:57
Data Information:
  • After 2026-05-05 06:03 UTC, this monitor no longer uses markdownified HTML/MDX. Comparisons across that boundary can therefore show more extensive diffs.

Translations

Create translation

post /audio/translations

Translates audio into English.

Returns

  • Translation object { text }

    • text: string
  • TranslationVerbose object { duration, language, text, segments }

    • duration: number

      The duration of the input audio.

    • language: string

      The language of the output translation (always english).

    • text: string

      The translated text.

    • segments: optional array of TranscriptionSegment

      Segments of the translated text and their corresponding details.

      • id: number

        Unique identifier of the segment.

      • avg_logprob: number

        Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

      • compression_ratio: number

        Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

      • end: number

        End time of the segment in seconds.

      • no_speech_prob: number

        Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

      • seek: number

        Seek offset of the segment.

      • start: number

        Start time of the segment in seconds.

      • temperature: number

        Temperature parameter used for generating the segment.

      • text: string

        Text content of the segment.

      • tokens: array of number

        Array of token IDs for the text content.

Example

curl https://api.openai.com/v1/audio/translations \
    -H 'Content-Type: multipart/form-data' \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -F 'file=@/path/to/file' \
    -F model=whisper-1

Response

{
  "text": "text"
}

Example

curl https://api.openai.com/v1/audio/translations \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: multipart/form-data" \
  -F file="@/path/to/file/german.m4a" \
  -F model="whisper-1"

Response

{
  "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}

Domain Types

Translation

  • Translation object { text }

    • text: string

Translation Verbose

  • TranslationVerbose object { duration, language, text, segments }

    • duration: number

      The duration of the input audio.

    • language: string

      The language of the output translation (always english).

    • text: string

      The translated text.

    • segments: optional array of TranscriptionSegment

      Segments of the translated text and their corresponding details.

      • id: number

        Unique identifier of the segment.

      • avg_logprob: number

        Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

      • compression_ratio: number

        Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

      • end: number

        End time of the segment in seconds.

      • no_speech_prob: number

        Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

      • seek: number

        Seek offset of the segment.

      • start: number

        Start time of the segment in seconds.

      • temperature: number

        Temperature parameter used for generating the segment.

      • text: string

        Text content of the segment.

      • tokens: array of number

        Array of token IDs for the text content.

Translation Create Response

  • TranslationCreateResponse = Translation or TranslationVerbose

    • Translation object { text }

      • text: string
    • TranslationVerbose object { duration, language, text, segments }

      • duration: number

        The duration of the input audio.

      • language: string

        The language of the output translation (always english).

      • text: string

        The translated text.

      • segments: optional array of TranscriptionSegment

        Segments of the translated text and their corresponding details.

        • id: number

          Unique identifier of the segment.

        • avg_logprob: number

          Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.

        • compression_ratio: number

          Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.

        • end: number

          End time of the segment in seconds.

        • no_speech_prob: number

          Probability of no speech in the segment. If the value is higher than 1.0 and the avg_logprob is below -1, consider this segment silent.

        • seek: number

          Seek offset of the segment.

        • start: number

          Start time of the segment in seconds.

        • temperature: number

          Temperature parameter used for generating the segment.

        • text: string

          Text content of the segment.

        • tokens: array of number

          Array of token IDs for the text content.