Translations
Create translation
client.Audio.Translations.New(ctx, body) (*Translation, error)
post /audio/translations
Translates audio into English.
Parameters
-
body AudioTranslationNewParams-
File param.Field[Reader]The audio file object (not file name) translate, in one of these formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm.
-
Model param.Field[AudioModel]ID of the model to use. Only
whisper-1(which is powered by our open source Whisper V2 model) is currently available.-
string -
type AudioModel string-
const AudioModelWhisper1 AudioModel = "whisper-1" -
const AudioModelGPT4oTranscribe AudioModel = "gpt-4o-transcribe" -
const AudioModelGPT4oMiniTranscribe AudioModel = "gpt-4o-mini-transcribe" -
const AudioModelGPT4oMiniTranscribe2025_12_15 AudioModel = "gpt-4o-mini-transcribe-2025-12-15" -
const AudioModelGPT4oTranscribeDiarize AudioModel = "gpt-4o-transcribe-diarize"
-
-
-
Prompt param.Field[string]An optional text to guide the model's style or continue a previous audio segment. The prompt should be in English.
-
ResponseFormat param.Field[AudioTranslationNewParamsResponseFormat]The format of the output, in one of these options:
json,text,srt,verbose_json, orvtt.-
const AudioTranslationNewParamsResponseFormatJSON AudioTranslationNewParamsResponseFormat = "json" -
const AudioTranslationNewParamsResponseFormatText AudioTranslationNewParamsResponseFormat = "text" -
const AudioTranslationNewParamsResponseFormatSRT AudioTranslationNewParamsResponseFormat = "srt" -
const AudioTranslationNewParamsResponseFormatVerboseJSON AudioTranslationNewParamsResponseFormat = "verbose_json" -
const AudioTranslationNewParamsResponseFormatVTT AudioTranslationNewParamsResponseFormat = "vtt"
-
-
Temperature param.Field[float64]The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.
-
Returns
-
type AudioTranslationNewResponse interface{…}-
type Translation struct{…}Text string
-
Example
package main
import (
"bytes"
"context"
"fmt"
"io"
"github.com/openai/openai-go"
"github.com/openai/openai-go/option"
)
func main() {
client := openai.NewClient(
option.WithAPIKey("My API Key"),
)
translation, err := client.Audio.Translations.New(context.TODO(), openai.AudioTranslationNewParams{
File: io.Reader(bytes.NewBuffer([]byte("Example data"))),
Model: openai.AudioModelWhisper1,
})
if err != nil {
panic(err.Error())
}
fmt.Printf("%+v\n", translation)
}
Response
{
"text": "text"
}
Domain Types
Translation
-
type Translation struct{…}Text string
Translation Verbose
-
type TranslationVerbose struct{…}-
Duration float64The duration of the input audio.
-
Language stringThe language of the output translation (always
english). -
Text stringThe translated text.
-
Segments []TranscriptionSegmentSegments of the translated text and their corresponding details.
-
ID int64Unique identifier of the segment.
-
AvgLogprob float64Average logprob of the segment. If the value is lower than -1, consider the logprobs failed.
-
CompressionRatio float64Compression ratio of the segment. If the value is greater than 2.4, consider the compression failed.
-
End float64End time of the segment in seconds.
-
NoSpeechProb float64Probability of no speech in the segment. If the value is higher than 1.0 and the
avg_logprobis below -1, consider this segment silent. -
Seek int64Seek offset of the segment.
-
Start float64Start time of the segment in seconds.
-
Temperature float64Temperature parameter used for generating the segment.
-
Text stringText content of the segment.
-
Tokens []int64Array of token IDs for the text content.
-
-