Skip to main content
Generate a voice from a text prompt without source audio. The flow has two steps:
  1. Generate a temporary preview. Use the live streaming endpoint for the fastest first audio, or the synchronous endpoint when you want multiple candidates or an encoded response.
  2. Pick the preview you like and call POST /v1/voice-previews/{generated_voice_id}/save to persist a real voice.
Choose model_id from List models and check languages[].language_id for the target language. Preview text can use Audio tags. For guidance on writing voice_description and preview text, see Voice Design prompting.

Stream one preview as it is generated

POST /v1/voice-previews/design/stream returns raw PCM as soon as the first audio chunk is available. The format is fixed to signed 16-bit little-endian PCM at 24 kHz, mono. The endpoint always creates one candidate, so it does not accept preview_count or output_format. The generated-voice-id response header identifies the temporary preview. Wait for a clean end of the response before calling the save endpoint: clean EOF means Breeze has persisted the completed WAV asset. If the connection fails or you stop reading early, discard the partial audio and do not retry by automatically switching to the synchronous endpoint.

Generate synchronous previews

Breeze bills by successful preview count times the billable units of the preview text (Chinese, Japanese and Korean characters count as two units each; see Pricing). voice_description and text each accept up to 500 characters. Set preview_count to 3 when you want three candidates in one request. Use guidance_scale to tune how strongly generation follows the voice description and preview script. Accepted values range from 1.0 to 10.0. When you pass it, that exact value is used. When you omit it, Breeze picks a random value between 1.0 and 10.0 for the request.

Download a completed preview

The synchronous response embeds preview audio in audio_base_64. You can also download any completed preview with GET /v1/voice-previews/{generated_voice_id}/stream. Despite its legacy path name, this GET endpoint downloads an already completed preview; it does not expose generation-time chunks.

Persist the voice

Pass the preview’s language when saving. See supported language codes. language_code is required so the saved voice always carries explicit reference-language metadata.
The saved voice works with POST /v1/text-to-speech/{voice_id} and appears on the voices page in the console.

Continue building

Voice Design prompting

Describe a reusable voice identity and write a representative preview script.

Stream a design preview

Start receiving one fixed-format PCM candidate while it is generated.

Create design previews

Inspect voice_description, text, preview_count, and guidance_scale with SDK examples.

Save voice preview

Persist a generated_voice_id as a reusable voice with a name and description.

CLI voice design

Design, audition, and save voice previews from the terminal with breeze voice design.

Voices

Understand designed voices, saved voices, voice IDs, and persisted voice settings.

Text to speech

Use the saved voice_id to generate sync, async, or streaming speech.

Pricing

Estimate preview cost, generation concurrency, and shared Studio/API credits.