Skip to main content
Breeze applies plan-based concurrent generation limits to protect long-running synthesis capacity. Standard API requests are not currently throttled by a published request-per-minute quota.

Concurrent generations

Concurrent generations are synthesis jobs that are admitted or running. Studio and Developer API jobs share the same per-user pool. Text to speech, streaming text to speech, async text to speech, voice design, and voice cloning all count toward these limits. Voice design counts each requested preview as one generation. When preview_count is omitted, Breeze generates one preview. When concurrent generation capacity is full, the API returns 429 GENERATION_CONCURRENCY_EXCEEDED with Retry-After. When Breeze shared generation capacity is temporarily unavailable, the API returns 503 GENERATION_CAPACITY_EXCEEDED.

Realtime sessions

Realtime text-to-speech WebSocket sessions have their own limits on top of the generation pool:
  • Physical connection attempts share an account-wide token bucket across all API keys and servers: 30 attempts per minute, with up to 10 immediately available burst slots. Successful and failed admitted attempts consume slots. This applies to both direct API-key and short-lived session-credential connections.
  • Upstream setup failures (including invalid handshake responses) and upstream failures after session.ready are counted separately for each key. Three consecutive failures of either kind, with less than 10 minutes between failures, trigger a 15-second cooldown; further failures increase it exponentially, up to 5 minutes. A successful handshake resets only setup failures. Delivered audio resets failures after ready. Closing an older session does not reset newer failures. An exhausted key budget or account balance triggers a 60-second cooldown. Other keys retain their own cooldown state.
  • Five consecutive empty sessions ending in idle timeout, with less than one hour between them, trigger a 60-second cooldown; repeated empty sessions increase it up to 5 minutes. Sending audio resets this streak. Reuse an active connection and send heartbeat messages while waiting for the next turn; do not continuously reconnect an unused session.
  • A rate or cooldown rejection arrives as an error frame with code: "RATE_LIMITED", meta.retry_after_seconds, meta.retry_after_ms, and close code 1008. Wait at least that long before explicitly opening a new session. Rejected attempts do not extend an existing cooldown. Existing connections and in-flight turns are not terminated by a new connection’s rejection.
  • Concurrent-session protection can also reject a connection with RATE_LIMITED, meta.max_sessions_per_key, and close code 1008. Close an existing idle session before opening a replacement when this limit is reached.
  • A session stays open for at most 30 minutes (SESSION_EXPIRED, close code 1000) and times out when idle for inactivity_timeout_seconds — 30 seconds by default, configurable up to 180 (IDLE_TIMEOUT, close code 1001). Frames from the client and audio for an in-flight turn both count as activity.
  • Managed SDK and CLI conversations normally mark each physical WebSocket for turn-boundary rotation when it reaches 10 minutes and continue on a fresh connection. An active turn can delay the switch; this client-side lifecycle does not change the 30-minute raw WebSocket limit or consume an extra generation while the conversation is idle.
  • Each realtime turn occupies one concurrent generation from the plan pool above while it synthesizes. When the pool is full, turn.start receives a GENERATION_CONCURRENCY_EXCEEDED error frame; the session stays open, so retry the turn with backoff.
Event payloads, error semantics, and close codes are documented in the realtime guide.

Best practices

  • Use exponential backoff with jitter for 429 GENERATION_CONCURRENCY_EXCEEDED and transient 5xx responses.
  • Group text into natural requests instead of sending one word at a time.
  • Treat async text to speech as a background job and poll for completion. Async delivery does not bypass concurrent generation limits.
  • Cache responses when appropriate. Identical inputs are re-billed.
  • For segmented live dialogue, use a bounded application-side semaphore, cancel stale remainder segments, and honor Retry-After. The realtime character response guide provides Python and TypeScript examples.

Design around limits

Pricing

See how plan limits and shared Studio/API credits fit together.

Text to speech

Use async jobs for longer text and retry only the failed segments of a batch.

Streaming

Choose streaming for lower time-to-first-byte without bypassing generation concurrency.

Errors

Handle 429, 503, and 504 responses with the right retry behavior.