Skip to main content
Voice cloning is a two-step flow: upload one sample to generate a preview, then save it as a reusable voice. Breeze detects the spoken language and automatically writes a short preview script in that language. Provide the audio and optional name and description; custom scripts, performance instructions, and language hints are not accepted. Speaker consent is required; see voice consent. To create your own lines or cross-language speech, save the voice and use Text to Speech.

Sample requirements

  • Upload exactly one sample file.
  • Use a single speaker with no background music or noise.
  • Use MP3 or WAV.
  • Minimum duration: 3 seconds. There is no maximum.
  • Maximum file size: 5 MB.

Step 1: Create a preview

Breeze analyzes the first 60 seconds of the upload: it transcribes that stretch, detects its language, and keeps at most 30 seconds of it, ending on the last complete sentence that fits. Without usable punctuation the cut falls back to a natural pause, and failing that to the last whole word — it never lands mid-word. The excerpt starts just before the first word, keeping up to a second of leading silence, and is normalized to -18 LUFS so a quiet or hot recording does not bias the clone. The stored reference is that trimmed, normalized excerpt together with its transcript, not the original file. A long recording therefore clones as well as a tightly edited one: upload the whole take and let Breeze pick the cut. Put the audio you want cloned in the first minute.
The response contains a generated_voice_id, not a saved voice_id. The preview script is generated by Breeze in the recognized audio language, without performance instructions. SDKs also accept files with exactly one item. If the reference speech cannot be identified, upload a clearer sample in a supported language.

Step 2: Stream the preview

Step 3: Save the preview as a voice

Pass the reference transcript’s language when saving. See supported language codes. language_code is required so later cross-language synthesis can distinguish the voice’s reference language from the requested speech language.
Saving consumes a voice slot. Subsequent POST /v1/text-to-speech/{voice_id} calls use the stored reference sample and its transcript, not the preview audio.

Editing or deleting a saved voice

  • Update tags and the description with PATCH /v1/voices/{voice_id}.
  • Tune defaults with PATCH /v1/voices/{voice_id}/settings.
  • Remove the voice with DELETE /v1/voices/{voice_id}.

Continue building

Create clone preview

Upload one audio sample with an optional name and description.

Download completed voice preview

Listen to a generated_voice_id preview before deciding to save it.

Save voice preview

Persist the preview as a reusable voice; saving consumes a voice slot.

CLI voice clone

Clone from an audio sample in the terminal with breeze voice clone.

Voices

Learn how cloned voices, designed voices, public voices, and voice settings fit together.

Text to speech

Use the saved cloned voice for dialogue lines, previews, or production audio.

Use the cloned voice

Save the preview, then use the saved voice_id with Text to Speech to choose your script, instructions, and target language. The saved voice retains its original reference audio and transcript.

Names and reference analysis

An omitted, null, empty, or whitespace-only Clone name uses the audio filename stem, or Cloned Voice when unavailable. Saving without a name keeps the preview name. Automatic names are limited to 80 Unicode characters; explicit names longer than 80 characters are rejected. Other preview types still require a name when saving. A nonempty transcript in a supported main ASR language may proceed despite low confidence or mixed-language evidence. Unsupported languages and empty transcripts remain blocked. If denoising removes usable speech, analysis tries the original audio once within the same request budget; the selected audio and transcript stay together. Service errors remain retryable and are not audio-quality verdicts. File types, size, duration, model selection, and pricing are unchanged.