language_code is optional for HTTP, streaming, async, realtime, and Voice Design requests. When it is omitted, the field stays absent and the model handles the text directly. Omission does not produce a validation error and does not inherit the reference voice’s language. If you send language_code explicitly, it must be a supported ISO 639-1 two-letter code for the selected model. See supported language codes. When model_id is omitted, explicit language_code: en or zh selects breeze-tts-2; other or omitted languages select breeze-tts-2-multilingual. Use List models and its languages[].language_id values to discover valid model-language combinations. Breeze does not inspect the script to route models. An explicit model is never switched automatically: an incompatible language returns an error.
HTTP text-to-speech requests accept up to 1,000 characters by default. Accounts with an approved higher limit can send up to their configured limit, capped at 2,000 characters, for synchronous, asynchronous, and HTTP streaming requests. The limit applies to each request; longer text is rejected rather than truncated. Existing SDK and CLI calls use the account limit without additional parameters.
For non-streaming requests with explicit language_code: "zh", text longer than 500 characters is split into at most five segments, each no longer than 500 characters. Splits prefer punctuation and whitespace, with a character-boundary fallback for long unpunctuated text. The original text is preserved. Segments are synthesized within the account’s concurrency limit, then joined in order with exactly 0.25 seconds of added silence between segments and no added silence at the end. Sync delivery returns the combined audio; async delivery produces one combined job result. Billing uses the original text once. The overall request character limit still applies. Realtime does not use this segmentation.
Japanese, Korean, automatic-language segmentation, and segmented HTTP streaming are available only where enabled. Automatic-language segmentation requires CJK characters to make up at least 80% of letters and digits; it does not change the requested model or language settings. Segmented HTTP streaming starts with the first segment and delivers subsequent segments in order. If generation fails after audio has started, the stream terminates with an error; discard the incomplete output. Splitting reduces long-input risk but does not guarantee that every word is spoken.
Audio tags
Audio tags add vocal actions and non-speech sounds at specific points in the input. Place the exact tag directly in your text:breeze-tts-2-multilingual. See Audio tags for the exact tag spellings across all supported languages.
Single request
wav_48000 profile:
output_format also takes a bitrate for the lossy encodings, as in mp3_44100_192. See Output formats for the accepted rates and bitrates.
Batching multiple lines
For dialogue, send one request per line and concatenate the audio. This keeps responses small, allows parallel calls, and limits retries to failed segments. Use async jobs or HTTP streaming when you already have complete text for each line. Use realtime WebSocket when your product is generating and speaking turns during a live conversation. For live characters that split one reply into a preamble, first sentence, and remainder, follow the realtime character response guide. It covers bounded concurrency, stale-segment cancellation, andRetry-After handling.
Realtime conversation
Realtime text to speech keeps one WebSocket connection open across multiple turns. It is optimized for voice agents, live chat, and barge-in flows where avoiding a new HTTP connection per turn matters.pcm_s16le, 24000 Hz, mono, 16-bit. It does not accept output_format; request mp3 or wav from the HTTP endpoints when you need encoded files.
Tuning expressive controls
For prompt-writing examples and a workflow for your own LLM, see Voice instruction prompting. Combineinstructions with voice_settings to control delivery per call. guidance_scale adjusts how strongly generation follows those instructions and the reference voice. The accepted range is 1.0 to 10.0.
Write instructions in the same language as the input text and language_code. The API, SDKs, CLI, and realtime WebSocket send instructions to the selected TTS model as authored; they do not translate them before inference. The Instruction Enhance endpoint also keeps the language you supplied.
Async jobs
Keep sync and streaming calls for short previews. Use async jobs for long text, reference-heavy voices, or batch production so the caller can poll status and download the audio after generation completes. Use the returnedgeneration_job_id with Get generation job, then Download generation job audio when ready.
Inspecting the result
The logs page shows inline playback, a download link, and Copy as cURL.Continue building
Voice instruction prompting
Direct the situation, intent, and delivery of each passage.
Convert text to speech
Inspect request fields, SDK examples, response bodies, and generated errors.
CLI text to speech
Generate, stream, and save speech from the terminal with
breeze tts.Streaming
Use the streaming endpoint when playback should begin before the full response is ready.
Realtime text to speech
Keep one WebSocket open across live conversation turns.
Realtime character response
Split live dialogue into responsive segments without retry storms.
Output formats
Pick an encoding for browser playback, post-processing, or low-latency audio pipelines.
Managing history
List previous generations, download audio, and delete test runs.
Speech speed
Setvoice_settings.speed in a TTS request to adjust speed without changing pitch. The range is 0.5–2.0. Omission always means 1.0, even if the voice has a stored speed setting. 0.5 approximately doubles the duration; 2.0 approximately halves it. Pauses scale together with speech.
voice_settings={"speed": 1.25} to convert, create_job, or stream. TypeScript: pass voiceSettings: { speed: 1.25 } in the request. CLI: use breeze tts "Hello from Breeze" --speed 1.25 --no-play.
Volume
Setvoice_settings.volume to apply a linear amplitude multiplier to the generated audio. The range is 0.01–2.0 (1%–200%), and omission means 1.0. This is a signal-level multiplier, not loudness normalization; values above 1.0 can clip peaks in recordings that are already near full scale.
voice_settings={"volume": 1.5}; TypeScript uses voiceSettings: { volume: 1.5 }; CLI uses breeze tts "Hello from Breeze" --volume 1.5 --no-play.
Word timing
Use Convert with timestamps for one complete JSON response containing base64 audio and the full word/token list. For incremental playback, use Stream speech with word timestamps. Both require a BreezeBlue TTS 2 model. The CLI supportsbreeze tts --with-timestamps for streaming and --no-stream --with-timestamps for complete delivery; both save a .timestamps.json file beside the decoded audio. The Speech timing guide includes HTTP, Python, TypeScript, and CLI examples for highlighting and seeking.
