Translations
Create translation
audio.translations.create(**kwargs) -> TranslationCreateResponse
post /audio/translations
Create translation
Parameters
-
file: StringThe audio file object (not file name) translate, in one of these formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm.
-
model: String | AudioModelID of the model to use. Only
whisper-1(which is powered by our open source Whisper V2 model) is currently available.-
String = String -
AudioModel = :"whisper-1" | :"gpt-4o-transcribe" | :"gpt-4o-mini-transcribe" | 2 more-
:"whisper-1" -
:"gpt-4o-transcribe" -
:"gpt-4o-mini-transcribe" -
:"gpt-4o-mini-transcribe-2025-12-15" -
:"gpt-4o-transcribe-diarize"
-
-
-
prompt: StringAn optional text to guide the model's style or continue a previous audio segment. The prompt should be in English.
-
response_format: :json | :text | :srt | 2 moreThe format of the output, in one of these options:
json,text,srt,verbose_json, orvtt.-
:json -
:text -
:srt -
:verbose_json -
:vtt
-
-
temperature: FloatThe sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.
Returns
-
TranslationCreateResponse = Translation | TranslationVerbose-
class Translationtext: String
-
class TranslationVerbose-
duration: FloatThe duration of the input audio.
-
language: StringThe language of the output translation (always
english). -
text: StringThe translated text.
-
segments: Array[TranscriptionSegment]Segments of the translated text and their corresponding details.
-
id: IntegerUnique identifier of the segment.
-
avg_logprob: FloatAverage logprob of the segment. If the value is lower than -1, consider the logprobs failed.
-
compression_ratio: FloatCompression ratio of the segment. If the value is greater than 2.4, consider the compression failed.
-
end_: FloatEnd time of the segment in seconds.
-
no_speech_prob: FloatProbability of no speech in the segment. If the value is higher than 1.0 and the
avg_logprobis below -1, consider this segment silent. -
seek: IntegerSeek offset of the segment.
-
start: FloatStart time of the segment in seconds.
-
temperature: FloatTemperature parameter used for generating the segment.
-
text: StringText content of the segment.
-
tokens: Array[Integer]Array of token IDs for the text content.
-
-
-
Example
require "openai"
openai = OpenAI::Client.new(api_key: "My API Key")
translation = openai.audio.translations.create(file: StringIO.new("Example data"), model: :"whisper-1")
puts(translation)
Response
{
"text": "text"
}
Domain Types
Translation
-
class Translationtext: String
Translation Verbose
-
class TranslationVerbose-
duration: FloatThe duration of the input audio.
-
language: StringThe language of the output translation (always
english). -
text: StringThe translated text.
-
segments: Array[TranscriptionSegment]Segments of the translated text and their corresponding details.
-
id: IntegerUnique identifier of the segment.
-
avg_logprob: FloatAverage logprob of the segment. If the value is lower than -1, consider the logprobs failed.
-
compression_ratio: FloatCompression ratio of the segment. If the value is greater than 2.4, consider the compression failed.
-
end_: FloatEnd time of the segment in seconds.
-
no_speech_prob: FloatProbability of no speech in the segment. If the value is higher than 1.0 and the
avg_logprobis below -1, consider this segment silent. -
seek: IntegerSeek offset of the segment.
-
start: FloatStart time of the segment in seconds.
-
temperature: FloatTemperature parameter used for generating the segment.
-
text: StringText content of the segment.
-
tokens: Array[Integer]Array of token IDs for the text content.
-
-
Translation Create Response
-
TranslationCreateResponse = Translation | TranslationVerbose-
class Translationtext: String
-
class TranslationVerbose-
duration: FloatThe duration of the input audio.
-
language: StringThe language of the output translation (always
english). -
text: StringThe translated text.
-
segments: Array[TranscriptionSegment]Segments of the translated text and their corresponding details.
-
id: IntegerUnique identifier of the segment.
-
avg_logprob: FloatAverage logprob of the segment. If the value is lower than -1, consider the logprobs failed.
-
compression_ratio: FloatCompression ratio of the segment. If the value is greater than 2.4, consider the compression failed.
-
end_: FloatEnd time of the segment in seconds.
-
no_speech_prob: FloatProbability of no speech in the segment. If the value is higher than 1.0 and the
avg_logprobis below -1, consider this segment silent. -
seek: IntegerSeek offset of the segment.
-
start: FloatStart time of the segment in seconds.
-
temperature: FloatTemperature parameter used for generating the segment.
-
text: StringText content of the segment.
-
tokens: Array[Integer]Array of token IDs for the text content.
-
-
-