Skip to main content
The breeze voice command family covers the full voice lifecycle: browse and audition voices, design new ones from a text description, clone from an audio sample, remix an existing voice, localize a voice into another language, tune generation settings, and manage what you keep. Personal aliases set in Voice Library → My Voices → Edit Voice appear in voice list, get, and search results for the same account. Search by the alias to find a saved public voice; its voice ID remains unchanged.

Browse and audition

Fetch the code contract before generating a form or editing metadata from an agent:
The JSON response contains language_codes, gender_codes, age_codes, tone_codes, tone_max_items, and accent_codes_by_language. Each supported language is present in the accent map; an empty array means --accent must be omitted or cleared for that language. voice search uses semantic and text relevance for public catalog voices and text or identifier matching for personal voices. voice list --search remains available for scripts that already use it and has the same search behavior. Without a search query, voice list defaults to creation time from newest to oldest (--sort created_at_unix --sort-direction desc). Use --sort name for alphabetical order. For the public catalog, --voice-type default --sort trend uses the shared Daily Trend snapshot; descending order means most popular first. Search results remain relevance-ranked, with Trend used only inside an equivalent relevance tier. --voice-type accepts all, default, or personal; results are paged with --page and --page-size. --origin accepts designed or cloned. Repeat --tag for exact AND filtering, and use --favorites-only to return only favorited voices. --primary-category is an exact filter. --accent american also includes English voices with a null accent. Use --language zh --accent-mode unmarked for Mandarin. --accent-mode accepts all (default) or unmarked; unmarked cannot be combined with --accent. Omitting both leaves accents unrestricted. Repeat --gender, --age, or --tone to match any supplied value within that metadata field. Trend pagination returns an opaque token; continue with --next-page-token <token> so the CLI keeps the original snapshot and filter set fixed across pages.

Design a voice

breeze voice design generates voice previews from a description. It consumes account credits. When a Breeze agent designs a voice, it explicitly requests standard Mandarin for Chinese and General American English for English unless you ask for another accent. The agent includes the pronunciation requirement in the description and auditions the preview before use. These are agent authoring defaults; the CLI does not automatically insert an accent into your description or change an existing voice or clone reference.
In an interactive terminal, the default single-preview command uses the live 24 kHz mono PCM endpoint and starts playback while the voice is still being generated. A clean end means the returned generated_voice_id is immediately ready for voice preview save. The CLI does not silently regenerate if a live stream fails after audio starts. --output json, --no-play, a non-PCM --format, and --preview-count greater than 1 use the synchronous preview endpoint so machine-readable output and multi-candidate behavior remain deterministic. Useful flags:
  • --preview-count — number of previews to generate (default 1).
  • --text — the line each preview speaks.
  • --format — wav or mp3; selecting either uses synchronous generation.
  • --play-all — play every preview in sequence; --no-play skips playback.
  • --guidance-scale — how strongly generation follows the description; omitted, Breeze varies it per preview.

Work with previews

Design, clone, remix, and localize produce previews identified by a generated_voice_id. Play or download them, then save the keeper as a voice:
voice preview save creates a saved voice in your account. --language is required and stores the reference language used by later synthesis. Add --tag <value> repeatedly to set its tags. Voice Metadata flags use stable lowercase codes. --tone is repeatable up to three times. --accent requires a matching --language when both are passed; English and Chinese have language-specific accent sets, while other languages do not accept an accent. When voice edit replaces reference audio with --file, pass --language for the new sample. Metadata-only edits preserve the saved language when the flag is omitted.

Clone a voice

breeze voice clone creates a voice clone preview from an audio sample you are authorized to use. It consumes account credits.
--file takes an MP3 or WAV of at least 3 seconds, up to 5 MB. Breeze identifies the reference language and automatically generates a preview script in that language. Clone previews accept no custom script, instructions, or language hints. Longer recordings are welcome: Breeze analyzes the first 60 seconds and stores at most 30 seconds of it — ending on a complete sentence and normalized to -18 LUFS — as the voice’s reference. Save the resulting preview with breeze voice preview save, the same as designed voices.

Remix a voice

Voice Remix changes an owned or official voice while keeping its language. The source stays unchanged; saving creates a separate private voice.
Each invocation generates one candidate and waits for it. Custom descriptions are enhanced automatically. An unchanged preset uses its curated instruction and script, translated into the source language when needed. Editing a preset with --prompt runs Enhance. breeze voice remix prepare --voice <voice_id> --prompt "..." --agent returns an instruction, matching script and a token valid for 24 hours without creating audio or reserving generation credits. To compare the Web Auto strengths, reuse that token for three generation commands at 4, 6 and 8. Each successful candidate costs 100 credits. Agent and JSON modes do not play audio automatically. A timeout includes the job ID; resume with voice remix status rather than creating another candidate. The status includes comparison_audio_url, comparison_text and source_voice. Audition the ready preview, then save it with voice preview save <generated_voice_id> --language <language_code>. Omitted name and description inherit the source values. See Voice Remix for the API and SDK workflow.

Localize a voice

breeze voice localize creates a preview of an existing voice speaking another language. It consumes account credits.
--voice is the source voice and --language is the target language code; both are required. --name sets the preview’s display name and defaults to the source voice name. Localization accepts no script, instructions, or reference file: Breeze writes the preview script in the target language. Localization runs as a generation job, so the command polls until the preview is ready, then plays it in interactive terminals. --wait-timeout bounds that wait, --format selects mp3 or wav for the downloaded preview, and --no-play skips playback for scripts and agents. Save the preview with breeze voice preview save <generated_voice_id> --name "..." --language <target_code>, which creates a separate voice and leaves the source voice unchanged.

Voice settings

Saved voices carry generation settings such as guidance strength:
voice settings edit accepts --guidance-scale, --stability, --similarity-boost, --style, --speed, and --use-speaker-boost; pass at least one. The saved --speed field is an inactive compatibility value (0.7–1.2). To change speech speed, use breeze tts --speed (0.5–2.0) for that request. Omitting the TTS flag means 1.0, regardless of saved settings.

Manage saved voices

voice edit --language corrects a saved voice’s language. Supported values are ar, cs, de, el, en, es, fi, fr, hi, id, it, ja, ko, nl, pl, pt, ro, ru, th, tr, uk, vi, and zh. voice edit accepts --gender, --age, repeatable --tone, and --accent. Use --clear-gender, --clear-age, --clear-tone, or --clear-accent to remove saved metadata. A set flag and its matching clear flag are mutually exclusive. Changing only --language automatically clears an existing accent that is invalid for the new language. voice tags set replaces the complete effective array. For a favorited public Voice it creates a personal override; voice tags clear resumes inheriting the owner’s tags. Unfavoriting keeps that override available if the Voice is favorited again. voice delete removes the voice from your Breeze account. It asks for confirmation in interactive terminals; scripts and agents must pass --yes or the command fails with a usage error. For an agent-safe lifecycle, use metadata-options, list or search, get, preview save, edit, and delete --yes with --agent. Read IDs and Metadata codes from JSON fields; do not parse the human-readable table output.

Prompt-writing guides

Voice Design prompting

Describe a reusable voice identity and write a representative preview script.

Designing a voice

Generate previews, audition them, and save a reusable voice.
For breeze voice clone, --name defaults to the filename stem or Cloned Voice. When saving, Clone previews default to their preview name and Remix previews default to the source voice’s name and description. Other preview types require --name.