breeze voice command family covers the full voice lifecycle: browse and audition voices, design new ones from a text description, clone from an audio sample, remix an existing voice, localize a voice into another language, tune generation settings, and manage what you keep.
Personal aliases set in Voice Library → My Voices → Edit Voice appear in
voice list, get, and search results for the same account. Search by the alias
to find a saved public voice; its voice ID remains unchanged.
Browse and audition
language_codes, gender_codes, age_codes,
tone_codes, tone_max_items, and accent_codes_by_language. Each supported
language is present in the accent map; an empty array means --accent must be
omitted or cleared for that language.
voice search uses semantic and text relevance for public catalog voices and
text or identifier matching for personal voices. voice list --search remains
available for scripts that already use it and has the same search behavior.
Without a search query, voice list defaults to creation time from newest to
oldest (--sort created_at_unix --sort-direction desc). Use --sort name for
alphabetical order. For the public catalog, --voice-type default --sort trend
uses the shared Daily Trend snapshot; descending order means most popular
first. Search results remain relevance-ranked, with Trend used only inside an
equivalent relevance tier.
--voice-type accepts all, default, or personal; results are paged with --page and --page-size.
--origin accepts designed or cloned. Repeat --tag for exact AND
filtering, and use --favorites-only to return only favorited voices.
--primary-category is an exact filter. --accent american also includes English
voices with a null accent. Use --language zh --accent-mode unmarked for Mandarin.
--accent-mode accepts all (default) or unmarked; unmarked cannot be combined
with --accent. Omitting both leaves accents unrestricted. Repeat --gender,
--age, or --tone to match any supplied value within that metadata field.
Trend pagination returns an opaque token; continue with
--next-page-token <token> so the CLI keeps the original snapshot and filter
set fixed across pages.
Design a voice
breeze voice design generates voice previews from a description. It consumes account credits.
When a Breeze agent designs a voice, it explicitly requests standard Mandarin for Chinese and General American English for English unless you ask for another accent. The agent includes the pronunciation requirement in the description and auditions the preview before use. These are agent authoring defaults; the CLI does not automatically insert an accent into your description or change an existing voice or clone reference.
generated_voice_id is immediately
ready for voice preview save. The CLI does not silently regenerate if a live
stream fails after audio starts.
--output json, --no-play, a non-PCM --format, and
--preview-count greater than 1 use the synchronous preview endpoint so
machine-readable output and multi-candidate behavior remain deterministic.
Useful flags:
--preview-count— number of previews to generate (default 1).--text— the line each preview speaks.--format—wavormp3; selecting either uses synchronous generation.--play-all— play every preview in sequence;--no-playskips playback.--guidance-scale— how strongly generation follows the description; omitted, Breeze varies it per preview.
Work with previews
Design, clone, remix, and localize produce previews identified by agenerated_voice_id. Play or download them, then save the keeper as a voice:
voice preview save creates a saved voice in your account. --language is
required and stores the reference language used by later synthesis. Add
--tag <value> repeatedly to set its tags.
Voice Metadata flags use stable lowercase codes. --tone is repeatable up to
three times. --accent requires a matching --language when both are passed;
English and Chinese have language-specific accent sets, while other languages
do not accept an accent.
When voice edit replaces reference audio with --file, pass --language for
the new sample. Metadata-only edits preserve the saved language when the flag
is omitted.
Clone a voice
breeze voice clone creates a voice clone preview from an audio sample you are authorized to use. It consumes account credits.
--file takes an MP3 or WAV of at least 3 seconds, up to 5 MB. Breeze identifies the reference language and automatically generates a preview script in that language. Clone previews accept no custom script, instructions, or language hints. Longer recordings are welcome: Breeze analyzes the first 60 seconds and stores at most 30 seconds of it — ending on a complete sentence and normalized to -18 LUFS — as the voice’s reference.
Save the resulting preview with breeze voice preview save, the same as designed voices.
Remix a voice
Voice Remix changes an owned or official voice while keeping its language. The source stays unchanged; saving creates a separate private voice.--prompt runs Enhance.
breeze voice remix prepare --voice <voice_id> --prompt "..." --agent returns an instruction, matching script and a token valid for 24 hours without creating audio or reserving generation credits. To compare the Web Auto strengths, reuse that token for three generation commands at 4, 6 and 8. Each successful candidate costs 100 credits.
Agent and JSON modes do not play audio automatically. A timeout includes the job ID; resume with voice remix status rather than creating another candidate. The status includes comparison_audio_url, comparison_text and source_voice. Audition the ready preview, then save it with voice preview save <generated_voice_id> --language <language_code>. Omitted name and description inherit the source values.
See Voice Remix for the API and SDK workflow.
Localize a voice
breeze voice localize creates a preview of an existing voice speaking another language. It consumes account credits.
--voice is the source voice and --language is the target language code; both
are required. --name sets the preview’s display name and defaults to the
source voice name. Localization accepts no script, instructions, or reference
file: Breeze writes the preview script in the target language.
Localization runs as a generation job, so the command polls until the preview is
ready, then plays it in interactive terminals. --wait-timeout bounds that wait,
--format selects mp3 or wav for the downloaded preview, and --no-play
skips playback for scripts and agents.
Save the preview with breeze voice preview save <generated_voice_id> --name "..." --language <target_code>, which creates a separate voice and leaves the source voice unchanged.
Voice settings
Saved voices carry generation settings such as guidance strength:voice settings edit accepts --guidance-scale, --stability, --similarity-boost, --style, --speed, and --use-speaker-boost; pass at least one.
The saved --speed field is an inactive compatibility value (0.7–1.2). To change speech speed, use breeze tts --speed (0.5–2.0) for that request. Omitting the TTS flag means 1.0, regardless of saved settings.
Manage saved voices
voice edit --language corrects a saved voice’s language. Supported values are ar, cs, de, el, en, es, fi, fr, hi, id, it, ja, ko, nl, pl, pt, ro, ru, th, tr, uk, vi, and zh.
voice edit accepts --gender, --age, repeatable --tone, and --accent.
Use --clear-gender, --clear-age, --clear-tone, or --clear-accent to
remove saved metadata. A set flag and its matching clear flag are mutually
exclusive. Changing only --language automatically clears an existing accent
that is invalid for the new language.
voice tags set replaces the complete effective array. For a favorited public
Voice it creates a personal override; voice tags clear resumes inheriting the
owner’s tags. Unfavoriting keeps that override available if the Voice is
favorited again.
voice delete removes the voice from your Breeze account. It asks for confirmation in interactive terminals; scripts and agents must pass --yes or the command fails with a usage error.
For an agent-safe lifecycle, use metadata-options, list or search, get,
preview save, edit, and delete --yes with --agent. Read IDs and Metadata
codes from JSON fields; do not parse the human-readable table output.
Prompt-writing guides
Voice Design prompting
Describe a reusable voice identity and write a representative preview script.
Designing a voice
Generate previews, audition them, and save a reusable voice.
breeze voice clone, --name defaults to the filename stem or Cloned Voice. When saving, Clone previews default to their preview name and Remix previews default to the source voice’s name and description. Other preview types require --name.
