Concurrent generations
Concurrent generations are synthesis jobs that are admitted or running. Studio and Developer API jobs share the same per-user pool. Text to speech, streaming text to speech, async text to speech, voice design, and voice cloning all count toward these limits. Voice design counts each requested preview as one generation. Whenpreview_count is omitted, Breeze generates one preview.
When concurrent generation capacity is full, the API returns
429 GENERATION_CONCURRENCY_EXCEEDED with Retry-After. When Breeze shared generation capacity is temporarily unavailable, the API returns 503 GENERATION_CAPACITY_EXCEEDED.
Realtime sessions
Realtime text-to-speech WebSocket sessions have their own limits on top of the generation pool:- Physical connection attempts share an account-wide token bucket across all API keys and servers: 30 attempts per minute, with up to 10 immediately available burst slots. Successful and failed admitted attempts consume slots. This applies to both direct API-key and short-lived session-credential connections.
- Upstream setup failures (including invalid handshake responses) and upstream failures after
session.readyare counted separately for each key. Three consecutive failures of either kind, with less than 10 minutes between failures, trigger a 15-second cooldown; further failures increase it exponentially, up to 5 minutes. A successful handshake resets only setup failures. Delivered audio resets failures after ready. Closing an older session does not reset newer failures. An exhausted key budget or account balance triggers a 60-second cooldown. Other keys retain their own cooldown state. - Five consecutive empty sessions ending in idle timeout, with less than one hour between them, trigger a 60-second cooldown; repeated empty sessions increase it up to 5 minutes. Sending audio resets this streak. Reuse an active connection and send heartbeat messages while waiting for the next turn; do not continuously reconnect an unused session.
- A rate or cooldown rejection arrives as an
errorframe withcode: "RATE_LIMITED",meta.retry_after_seconds,meta.retry_after_ms, and close code 1008. Wait at least that long before explicitly opening a new session. Rejected attempts do not extend an existing cooldown. Existing connections and in-flight turns are not terminated by a new connection’s rejection. - Concurrent-session protection can also reject a connection with
RATE_LIMITED,meta.max_sessions_per_key, and close code 1008. Close an existing idle session before opening a replacement when this limit is reached. - A session stays open for at most 30 minutes (
SESSION_EXPIRED, close code 1000) and times out when idle forinactivity_timeout_seconds— 30 seconds by default, configurable up to 180 (IDLE_TIMEOUT, close code 1001). Frames from the client and audio for an in-flight turn both count as activity. - Managed SDK and CLI conversations normally mark each physical WebSocket for turn-boundary rotation when it reaches 10 minutes and continue on a fresh connection. An active turn can delay the switch; this client-side lifecycle does not change the 30-minute raw WebSocket limit or consume an extra generation while the conversation is idle.
- Each realtime turn occupies one concurrent generation from the plan pool above while it synthesizes. When the pool is full,
turn.startreceives aGENERATION_CONCURRENCY_EXCEEDEDerror frame; the session stays open, so retry the turn with backoff.
Best practices
- Use exponential backoff with jitter for
429 GENERATION_CONCURRENCY_EXCEEDEDand transient5xxresponses. - Group text into natural requests instead of sending one word at a time.
- Treat async text to speech as a background job and poll for completion. Async delivery does not bypass concurrent generation limits.
- Cache responses when appropriate. Identical inputs are re-billed.
- For segmented live dialogue, use a bounded application-side semaphore, cancel stale remainder segments, and honor
Retry-After. The realtime character response guide provides Python and TypeScript examples.
Design around limits
Pricing
See how plan limits and shared Studio/API credits fit together.
Text to speech
Use async jobs for longer text and retry only the failed segments of a batch.
Streaming
Choose streaming for lower time-to-first-byte without bypassing generation concurrency.
Errors
Handle
429, 503, and 504 responses with the right retry behavior.
