SpyBara
Go Premium

go/resources/audio/subresources/speech/index.md 2026-05-02 05:57 UTC to 2026-05-05 23:00 UTC

152 added, 0 removed.

2026
Wed 27 06:42 Fri 22 06:33 Wed 20 06:35 Tue 19 06:34 Mon 18 22:01 Mon 11 18:00 Thu 7 21:57 Tue 5 23:00 Sat 2 05:57
Data Information:
  • After 2026-05-05 06:03 UTC, this monitor no longer uses markdownified HTML/MDX. Comparisons across that boundary can therefore show more extensive diffs.

Speech

Create speech

client.Audio.Speech.New(ctx, body) (*Response, error)

post /audio/speech

Generates audio from the input text.

Returns the audio file content, or a stream of audio events.

Parameters

  • body AudioSpeechNewParams

    • Input param.Field[string]

      The text to generate audio for. The maximum length is 4096 characters.

    • Model param.Field[SpeechModel]

      One of the available TTS models: tts-1, tts-1-hd, gpt-4o-mini-tts, or gpt-4o-mini-tts-2025-12-15.

      • string

      • type SpeechModel string

        • const SpeechModelTTS1 SpeechModel = "tts-1"

        • const SpeechModelTTS1HD SpeechModel = "tts-1-hd"

        • const SpeechModelGPT4oMiniTTS SpeechModel = "gpt-4o-mini-tts"

        • const SpeechModelGPT4oMiniTTS2025_12_15 SpeechModel = "gpt-4o-mini-tts-2025-12-15"

    • Voice param.Field[AudioSpeechNewParamsVoiceUnion]

      The voice to use when generating the audio. Supported built-in voices are alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar. You may also provide a custom voice object with an id, for example { "id": "voice_1234" }. Previews of the voices are available in the Text to speech guide.

      • string

      • type AudioSpeechNewParamsVoiceString string

        • const AudioSpeechNewParamsVoiceStringAlloy AudioSpeechNewParamsVoiceString = "alloy"

        • const AudioSpeechNewParamsVoiceStringAsh AudioSpeechNewParamsVoiceString = "ash"

        • const AudioSpeechNewParamsVoiceStringBallad AudioSpeechNewParamsVoiceString = "ballad"

        • const AudioSpeechNewParamsVoiceStringCoral AudioSpeechNewParamsVoiceString = "coral"

        • const AudioSpeechNewParamsVoiceStringEcho AudioSpeechNewParamsVoiceString = "echo"

        • const AudioSpeechNewParamsVoiceStringSage AudioSpeechNewParamsVoiceString = "sage"

        • const AudioSpeechNewParamsVoiceStringShimmer AudioSpeechNewParamsVoiceString = "shimmer"

        • const AudioSpeechNewParamsVoiceStringVerse AudioSpeechNewParamsVoiceString = "verse"

        • const AudioSpeechNewParamsVoiceStringMarin AudioSpeechNewParamsVoiceString = "marin"

        • const AudioSpeechNewParamsVoiceStringCedar AudioSpeechNewParamsVoiceString = "cedar"

      • type AudioSpeechNewParamsVoiceID struct{…}

        Custom voice reference.

        • ID string

          The custom voice ID, e.g. voice_1234.

    • Instructions param.Field[string]

      Control the voice of your generated audio with additional instructions. Does not work with tts-1 or tts-1-hd.

    • ResponseFormat param.Field[AudioSpeechNewParamsResponseFormat]

      The format to audio in. Supported formats are mp3, opus, aac, flac, wav, and pcm.

      • const AudioSpeechNewParamsResponseFormatMP3 AudioSpeechNewParamsResponseFormat = "mp3"

      • const AudioSpeechNewParamsResponseFormatOpus AudioSpeechNewParamsResponseFormat = "opus"

      • const AudioSpeechNewParamsResponseFormatAAC AudioSpeechNewParamsResponseFormat = "aac"

      • const AudioSpeechNewParamsResponseFormatFLAC AudioSpeechNewParamsResponseFormat = "flac"

      • const AudioSpeechNewParamsResponseFormatWAV AudioSpeechNewParamsResponseFormat = "wav"

      • const AudioSpeechNewParamsResponseFormatPCM AudioSpeechNewParamsResponseFormat = "pcm"

    • Speed param.Field[float64]

      The speed of the generated audio. Select a value from 0.25 to 4.0. 1.0 is the default.

    • StreamFormat param.Field[AudioSpeechNewParamsStreamFormat]

      The format to stream the audio in. Supported formats are sse and audio. sse is not supported for tts-1 or tts-1-hd.

      • const AudioSpeechNewParamsStreamFormatSSE AudioSpeechNewParamsStreamFormat = "sse"

      • const AudioSpeechNewParamsStreamFormatAudio AudioSpeechNewParamsStreamFormat = "audio"

Returns

  • type AudioSpeechNewResponse interface{…}

Example

package main

import (
  "context"
  "fmt"

  "github.com/openai/openai-go"
  "github.com/openai/openai-go/option"
)

func main() {
  client := openai.NewClient(
    option.WithAPIKey("My API Key"),
  )
  speech, err := client.Audio.Speech.New(context.TODO(), openai.AudioSpeechNewParams{
    Input: "input",
    Model: openai.SpeechModelTTS1,
    Voice: openai.AudioSpeechNewParamsVoiceUnion{
      OfString: openai.String("string"),
    },
  })
  if err != nil {
    panic(err.Error())
  }
  fmt.Printf("%+v\n", speech)
}

Domain Types

Speech Model

  • type SpeechModel string

    • const SpeechModelTTS1 SpeechModel = "tts-1"

    • const SpeechModelTTS1HD SpeechModel = "tts-1-hd"

    • const SpeechModelGPT4oMiniTTS SpeechModel = "gpt-4o-mini-tts"

    • const SpeechModelGPT4oMiniTTS2025_12_15 SpeechModel = "gpt-4o-mini-tts-2025-12-15"