voice_id.
Sources
Voices come from these sources:- Public catalog: curated voices that ship with Breeze. List them with
GET /v1/voices. - Designed voices: created from a text prompt with
POST /v1/voice-previews/designand finalized withPOST /v1/voice-previews/{generated_voice_id}/save. - Cloned voices: created from audio samples with
POST /v1/voice-previews/clone, previewed viaGET /v1/voice-previews/{generated_voice_id}/stream, and finalized withPOST /v1/voice-previews/{generated_voice_id}/save. - Localized voices: created from an existing voice with
POST /v1/voice-previews/localize, polled withGET /v1/voice-previews/localize/{generation_job_id}, and finalized withPOST /v1/voice-previews/{generated_voice_id}/save. The source voice is unchanged; see Voice Localize.
voice_type="default" is the official Breeze catalog. voice_type="personal" is user-saved voices. Categories are premade, generated, and cloned.
Daily Trend browsing
Usevoice_type=default&sort=trend to browse the daily ranking. Eligible public voices absent from the snapshot appear after ranked voices, ordered by voice ID. They can be found through metadata filters immediately, including voices made public after creation. sort_direction=asc reverses the ranked portion; supplemental voices remain last.
Continue with the returned next_page_token. It pins the daily snapshot and the initially eligible supplemental candidates for up to two hours. Newly published voices appear in a fresh query; removed or private voices disappear from subsequent pages. Expired tokens require a fresh first-page request. Text search remains relevance-ranked.
Voice Metadata
Voice responses always includelanguage_code, gender, age, tone, and
accent. These fields contain stable codes rather than localized display text, so you
can localize them in your own interface without changing stored values.
Call Metadata options before rendering a form or validating an
automated edit. The response contains the complete language_codes,
gender_codes, age_codes, and tone_codes arrays, the tone_max_items
limit, and accent_codes_by_language keyed by every supported language. An
empty accent array means the language requires accent: null.
genderismale,female,neutral, ornull.ageischild,young,middle_aged,old, ornull.toneis an array containing up to three distinct codes:warm,calm,bright,gentle,energetic,authoritative,sincere,weary,precise,refined,urgent,friendly,articulate,steady,playful,compassionate,reflective,measured,rhythmic, orpassionate. An unclassified voice returns an empty array.accentis a language-specific code ornull.
en) supports american, british, scottish, irish,
australian, canadian, us_southern, us_new_york, indian,
south_african, russian, japanese, korean, and chinese.
Chinese (zh) uses null for Mandarin and supports mandarin_northeastern,
mandarin_shaanxi, mandarin_shanghai, mandarin_sichuan, mandarin_henan,
mandarin_beijing, mandarin_taiwan, and cantonese. English null accents are
treated as American English for display and filtering. Other languages require
accent: null without an inferred accent. The API preserves null in responses.
Set these fields while saving a preview or update them with
PATCH /v1/voices/{voice_id}. For edits, omitting a field preserves its value;
an empty gender, age, or accent form field clears that value, and one
empty tone form field clears the array. If you change language_code without
setting accent, Breeze clears an existing accent that is invalid for the new
language. An explicitly incompatible language and accent returns
VALIDATION_ERROR.
The complete language list and the meaning of language_code across Voice and
text-to-speech requests are documented in Multilingual audio.
Category codes
Pass one of these exactprimary_category_code values to
List voices to filter the public catalog.
Categories describe use cases; they are separate from voice origin, language,
and Voice Metadata. This filter does not assign a category when saving or
editing a voice.
The List voices parameter schema exposes this same enum in
OpenAPI. Metadata options supplies language, gender, age,
tone, and accent codes; it does not return category codes.
Tags and favorites
Every Voice response exposes one effectivetags array. Breeze stores each tag
with the spelling you send: it normalizes Unicode to NFC, trims surrounding
whitespace, removes empty values, and accepts up to 20 tags of 128 characters
each. Comparison is case-insensitive: Sci-Fi and sci-fi are the same tag, so
the first spelling in an array wins and later case variants are dropped.
Tags are full-value strings. Strings such as 三体, 小说=三体, and
小说:科幻:三体 are all valid. The = and : characters may receive
structured styling in Breeze interfaces, but do not change API matching
semantics.
The owner maintains the Voice defaults. A user who favorites a public Voice can
replace those defaults with a complete personal array through
PUT /v1/voices/{voice_id}/tags. Sending an empty array removes the personal
override and immediately resumes inheritance from the owner. Unfavoriting hides
the override; favoriting the Voice again restores it.
Use repeated tags query parameters for full-value, case-insensitive AND
filtering, and favorites_only=true to restrict GET /v1/voices to favorites.
Manage voices with the API
Use the resource endpoints in this order for a predictable CRUD workflow:
Python SDK users can call
client.voices.metadata_options(), while TypeScript
SDK users can call client.voices.metadataOptions(). The SDK methods map the
same operations to search, get, save_preview / savePreview, edit,
update_tags / updateTags, favorite, unfavorite, and delete without
changing the HTTP semantics.
Voice settings
Each voice stores defaultvoice_settings. Pass per-request settings in the request body. guidance_scale inherits the saved value when omitted. For HTTP TTS, speed is request-scoped and defaults to 1.0; it does not inherit a saved speed and is unavailable in Realtime.
guidance_scaleadjusts how strongly generation follows the prompt and reference voice. Accepted values range from1.0to10.0.
Browsing voices
Passsearch to GET /v1/voices to search every voice available to the
authenticated account. Public catalog results combine vector similarity with
exact, prefix, and substring name matches. Personal voices remain searchable
by name, description, or voice ID. Exact name matches rank first, and the API
falls back to text matching if semantic search is temporarily unavailable.
Saved public voices can have a personal alias. Set or clear the alias in
Voice Library → My Voices → Edit Voice. The alias appears as name in
authenticated voice responses and is searchable in the app and API. It is
private to your account; the public voice name, voice ID, and other users’
results stay the same. Clearing the alias restores the original display name.
Cache voice-list responses in the client for at least 60 seconds. Refresh on an
explicit user action or after your application creates, edits, favorites, or
deletes a voice; do not poll the catalog continuously while an application is
idle.
Pass language_code to restrict the search to voices whose saved preview audio
uses that supported two-letter language code. See supported language codes. The language filter is applied
before semantic and text ranking, so every returned voice matches the requested
audio language.
The public catalog also supports exact primary_category_code and language-aware accent
filters, repeatable gender and age filters with OR semantics inside each
field, and repeatable tone filters that match any supplied tone. Combine
fields to require every selected metadata group.
Pass tags repeatedly to require every listed effective tag. Tag values match
the full string ignoring case; filtering does not perform hierarchy or prefix
matching.
For a public-catalog browse experience, pass voice_type=default&sort=trend.
The default descending direction returns the most popular voices from one
immutable global daily snapshot. Follow next_page_token rather than changing
page; the token pins the snapshot and filters so midnight refreshes cannot
reorder an in-progress traversal. Searches remain relevance-first: Trend only
breaks ties inside an equivalent lexical or semantic relevance tier.
Use the voice library to browse the public catalog and preview samples.
Voice workflows
Text to speech
Use any saved or public
voice_id to generate one-shot, async, or streaming audio.Voice Design
Create a new voice from a text description and sample script.
Voice Clone
Create a reusable voice from a consented audio sample.
Voice Localize
Create a voice that speaks another language from a voice you already have.
List voices
Search and filter the public catalog and your saved voices by
voice_type and origin.Metadata options
Read every valid code and the language-to-accent cascade before saving or editing a voice.
CLI voices
Browse, audition, design, clone, and manage voices from the terminal.
Voice Library
Browse voices visually, preview samples, and copy voice IDs from Breeze Studio.

