Skip to main content
A voice is the speaker identity used to synthesize audio. Every text-to-speech request targets a specific voice via its voice_id.

Sources

Voices come from these sources:
  • Public catalog: curated voices that ship with Breeze. List them with GET /v1/voices.
  • Designed voices: created from a text prompt with POST /v1/voice-previews/design and finalized with POST /v1/voice-previews/{generated_voice_id}/save.
  • Cloned voices: created from audio samples with POST /v1/voice-previews/clone, previewed via GET /v1/voice-previews/{generated_voice_id}/stream, and finalized with POST /v1/voice-previews/{generated_voice_id}/save.
  • Localized voices: created from an existing voice with POST /v1/voice-previews/localize, polled with GET /v1/voice-previews/localize/{generation_job_id}, and finalized with POST /v1/voice-previews/{generated_voice_id}/save. The source voice is unchanged; see Voice Localize.
voice_type="default" is the official Breeze catalog. voice_type="personal" is user-saved voices. Categories are premade, generated, and cloned.

Daily Trend browsing

Use voice_type=default&sort=trend to browse the daily ranking. Eligible public voices absent from the snapshot appear after ranked voices, ordered by voice ID. They can be found through metadata filters immediately, including voices made public after creation. sort_direction=asc reverses the ranked portion; supplemental voices remain last. Continue with the returned next_page_token. It pins the daily snapshot and the initially eligible supplemental candidates for up to two hours. Newly published voices appear in a fresh query; removed or private voices disappear from subsequent pages. Expired tokens require a fresh first-page request. Text search remains relevance-ranked.

Voice Metadata

Voice responses always include language_code, gender, age, tone, and accent. These fields contain stable codes rather than localized display text, so you can localize them in your own interface without changing stored values. Call Metadata options before rendering a form or validating an automated edit. The response contains the complete language_codes, gender_codes, age_codes, and tone_codes arrays, the tone_max_items limit, and accent_codes_by_language keyed by every supported language. An empty accent array means the language requires accent: null.
  • gender is male, female, neutral, or null.
  • age is child, young, middle_aged, old, or null.
  • tone is an array containing up to three distinct codes: warm, calm, bright, gentle, energetic, authoritative, sincere, weary, precise, refined, urgent, friendly, articulate, steady, playful, compassionate, reflective, measured, rhythmic, or passionate. An unclassified voice returns an empty array.
  • accent is a language-specific code or null.
English (en) supports american, british, scottish, irish, australian, canadian, us_southern, us_new_york, indian, south_african, russian, japanese, korean, and chinese. Chinese (zh) uses null for Mandarin and supports mandarin_northeastern, mandarin_shaanxi, mandarin_shanghai, mandarin_sichuan, mandarin_henan, mandarin_beijing, mandarin_taiwan, and cantonese. English null accents are treated as American English for display and filtering. Other languages require accent: null without an inferred accent. The API preserves null in responses. Set these fields while saving a preview or update them with PATCH /v1/voices/{voice_id}. For edits, omitting a field preserves its value; an empty gender, age, or accent form field clears that value, and one empty tone form field clears the array. If you change language_code without setting accent, Breeze clears an existing accent that is invalid for the new language. An explicitly incompatible language and accent returns VALIDATION_ERROR. The complete language list and the meaning of language_code across Voice and text-to-speech requests are documented in Multilingual audio.

Category codes

Pass one of these exact primary_category_code values to List voices to filter the public catalog. Categories describe use cases; they are separate from voice origin, language, and Voice Metadata. This filter does not assign a category when saving or editing a voice. The List voices parameter schema exposes this same enum in OpenAPI. Metadata options supplies language, gender, age, tone, and accent codes; it does not return category codes.

Tags and favorites

Every Voice response exposes one effective tags array. Breeze stores each tag with the spelling you send: it normalizes Unicode to NFC, trims surrounding whitespace, removes empty values, and accepts up to 20 tags of 128 characters each. Comparison is case-insensitive: Sci-Fi and sci-fi are the same tag, so the first spelling in an array wins and later case variants are dropped. Tags are full-value strings. Strings such as 三体, 小说=三体, and 小说:科幻:三体 are all valid. The = and : characters may receive structured styling in Breeze interfaces, but do not change API matching semantics. The owner maintains the Voice defaults. A user who favorites a public Voice can replace those defaults with a complete personal array through PUT /v1/voices/{voice_id}/tags. Sending an empty array removes the personal override and immediately resumes inheritance from the owner. Unfavoriting hides the override; favoriting the Voice again restores it. Use repeated tags query parameters for full-value, case-insensitive AND filtering, and favorites_only=true to restrict GET /v1/voices to favorites.

Manage voices with the API

Use the resource endpoints in this order for a predictable CRUD workflow: Python SDK users can call client.voices.metadata_options(), while TypeScript SDK users can call client.voices.metadataOptions(). The SDK methods map the same operations to search, get, save_preview / savePreview, edit, update_tags / updateTags, favorite, unfavorite, and delete without changing the HTTP semantics.

Voice settings

Each voice stores default voice_settings. Pass per-request settings in the request body. guidance_scale inherits the saved value when omitted. For HTTP TTS, speed is request-scoped and defaults to 1.0; it does not inherit a saved speed and is unavailable in Realtime.
  • guidance_scale adjusts how strongly generation follows the prompt and reference voice. Accepted values range from 1.0 to 10.0.
Update a voice’s persisted defaults with Update voice settings. For preview generation, see Voice Design controls, including the behavior when guidance is omitted.

Browsing voices

Pass search to GET /v1/voices to search every voice available to the authenticated account. Public catalog results combine vector similarity with exact, prefix, and substring name matches. Personal voices remain searchable by name, description, or voice ID. Exact name matches rank first, and the API falls back to text matching if semantic search is temporarily unavailable. Saved public voices can have a personal alias. Set or clear the alias in Voice Library → My Voices → Edit Voice. The alias appears as name in authenticated voice responses and is searchable in the app and API. It is private to your account; the public voice name, voice ID, and other users’ results stay the same. Clearing the alias restores the original display name. Cache voice-list responses in the client for at least 60 seconds. Refresh on an explicit user action or after your application creates, edits, favorites, or deletes a voice; do not poll the catalog continuously while an application is idle. Pass language_code to restrict the search to voices whose saved preview audio uses that supported two-letter language code. See supported language codes. The language filter is applied before semantic and text ranking, so every returned voice matches the requested audio language. The public catalog also supports exact primary_category_code and language-aware accent filters, repeatable gender and age filters with OR semantics inside each field, and repeatable tone filters that match any supplied tone. Combine fields to require every selected metadata group. Pass tags repeatedly to require every listed effective tag. Tag values match the full string ignoring case; filtering does not perform hierarchy or prefix matching. For a public-catalog browse experience, pass voice_type=default&sort=trend. The default descending direction returns the most popular voices from one immutable global daily snapshot. Follow next_page_token rather than changing page; the token pins the snapshot and filters so midnight refreshes cannot reorder an in-progress traversal. Searches remain relevance-first: Trend only breaks ties inside an equivalent lexical or semantic relevance tier. Use the voice library to browse the public catalog and preview samples.

Voice workflows

Text to speech

Use any saved or public voice_id to generate one-shot, async, or streaming audio.

Voice Design

Create a new voice from a text description and sample script.

Voice Clone

Create a reusable voice from a consented audio sample.

Voice Localize

Create a voice that speaks another language from a voice you already have.

List voices

Search and filter the public catalog and your saved voices by voice_type and origin.

Metadata options

Read every valid code and the language-to-accent cascade before saving or editing a voice.

CLI voices

Browse, audition, design, clone, and manage voices from the terminal.

Voice Library

Browse voices visually, preview samples, and copy voice IDs from Breeze Studio.