Create a video
Visual variety and motion
For narrated image-led videos, the agent chooses images and shot lengths from the script’s events and the measured audio, consulting matching SRT cues when available. New events or scenes normally receive new source images; the same image can stay on screen while its narrative moment continues. There are no fixed image durations, source quotas or reuse intervals. Purposeful flashbacks, comparisons and recurring motifs are recorded inSTORYBOARD.md. Reframing or animating the same image does not make it a new source. Prepared audio defines timing; SRT is exported during rendering and is not required before planning.
compose, inspect and render expose visual_coverage, including exact source hashes, image-use counts, cumulative exposure and reuse gaps. HTML image dependencies are listed for review; references do not measure their visibility. The agent reviews every recurrence and resolves unrelated repetition before delivery. Numeric coverage warnings and suggested_min_image_sources are diagnostic hints, not instructions to cut a shot or add images. The agent reviews warnings against the story and records the decision. Use inspect --strict to fail on warnings.
Image shots default to a static hold and can use native push-in or single-axis pans. Custom HTML motion must produce the same frame when seeking to the same timestamp. Inspection compares sampled forward, repeated/backward and freshly loaded seeks; mismatches block rendering. Review consecutive frames too, since repeatable motion can still have abrupt speed or direction changes.
Sound effects and ambience
The agent considers sound design from your script and visible events. It can obtain suitable sound assets from supplied files, available host audio tools, licensed recordings or local procedural synthesis. Effects are optional: cuts do not automatically receive a whoosh, and narration alone is a valid result. The CLI mixes local assets; it does not include an AI sound generator. An optionalsound-design.json uses version: "video-lab.sound-design.v1", mode: "auto" and a cues array. Each cue records a unique id, kind (effect or ambience), project-local asset, audio start/end anchors, purpose and source. Optional gain, fades, source offset and explicit looping control its treatment. Narration-driven ducking defaults to enabled. The bundled Video Lab Skill includes the complete contract and examples.
If you ask for no sound effects, the agent skips effect/ambience generation and retrieval, sets the plan’s mode to off, and preserves narration:
--no-sound-effects overrides even stale or malformed plans. It is also available on preflight and batch; batch jobs accept no_sound_effects: true. --no-audio instead removes all audio. Without a plan, or with an empty cue list, existing projects keep narration alone.
The render report and lock record audio.sound_design with the decision, cue timing, source hashes and mix provenance. Enabled cues are mixed before final loudness mastering; raw TTS and narration remain unchanged. Sound edits require another render, reuse visual chunks and never generate new TTS. Review the exported soundtrack for speech clarity, event alignment, levels and clean fades/loops.
Credits and dependencies
Audio generation, rendering, preflight and actual batch runs require an authenticated Breeze account with at least 400 credits. Only newly generated TTS consumes credits. Rendering never implicitly calls TTS.render --no-audio produces a silent rehearsal using the prepared audio timeline and also requires 400 credits. Setup, composition, capture, validation, inspection, snapshots and batch dry runs do not require an account.
Rendering needs Node.js, FFmpeg, ffprobe and Playwright Chromium. Audio preparation needs no browser. Run doctor against the installed CLI runtime to check its dependencies; a Playwright installation in another project may not be available to that runtime.
Revise without starting over
Audio segments retain their job IDs and verified cache files. Changes to a sentence, voice or performance direction regenerate only affected speech. Pause changes reassemble the audio; visual changes reuse it.Project and output
Projects usevideo-lab.v2: manifest.json defines dimensions and frame rate, script.json defines narration and roles, and storyboard.json defines visual shots. Keep editable HTML in scenes/ and imagery, fonts and libraries in assets/.
Generated audio, compositions, clips and reports live under output/. Do not edit generated compositions directly. The final MP4, render report and lock record its source fingerprints and audio provenance.
Long videos render in resumable chunks of at most 300 frames. Temporary PNGs are removed unless --keep-frames is requested. --force replaces the final MP4 while reusing valid caches.
Video Lab preserves complete narration. It does not shorten speech to fit a shot, generate sound assets, render in the cloud or publish your video.
Unspecified pauses between script segments default to 0.6 seconds of added silence; the final segment defaults to zero and uses tail_seconds. Explicit pause_after_seconds, including zero, takes precedence. Automatic TTS chunks add 0.25 seconds after sentence endings or 0.6 seconds at blank-line paragraph boundaries, and no extra silence at forced mid-sentence splits. These are added durations, not measured speech-to-speech gaps: existing recording silence is preserved. Review the actual joins and adjust semantic segments for pacing, transitions and chapter changes. audio.json lists automatic chunk joins and their added silence. Run video audio again to reassemble with cached speech, then recompose after timing changes; existing prepared audio is not rewritten automatically.
Every successful render writes UTF-8 final.srt beside final.mp4 and returns final_srt; the report and lock include its path, SHA-256, cue count and timing method. Captions use the original narration text, wrapped into at most two 42-character lines per cue. Placement within each generated TTS piece is estimated by text length, not word-aligned; explicit inter-segment and recorded inter-piece pauses are excluded. Review timing before publication. The SRT is a separate file and is not burned into the video. Silent rehearsals also export the prepared narration transcript. Existing prepared audio remains usable.
Narration direction and guidance
New projects explicitly start the narrator atguidance_scale: 1 and include role and segment delivery instructions. For continuous narration, start at 1 and audition 1–2 for stronger expression; the supported range remains 1–10. This does not guarantee identical timbre across independent generations.
Set stable speaker direction in voices.<role>.instructions, then describe emphasis, pace and emotional changes in each segment’s instructions. These instructions append to the role instructions. A segment can cover several connected sentences; sentence-level performance planning does not require separate generation for every sentence.
An optional segment guidance_scale overrides the role value. Omitted values in existing projects retain server inheritance: an applicable preset, saved voice settings, or the server fallback of 1. Changes to effective request instructions or guidance regenerate affected pieces; unchanged audio is reused.
output/audio/audio.json includes advisory warnings for missing segment direction, inherited guidance and guidance above 2. Each piece’s request_parameters records merged instructions, language, requested guidance and guidance_scale_source (segment, voice, or server_default). The server_default source means the field was omitted, not that the server necessarily used 1. effective_guidance_scale is null and its source is unverified until server execution is verified; historical cached pieces without snapshots are reported with null request parameters. Technical success does not certify timbre consistency or delivery quality: listen to representative passages and the complete assembled narration before publication.
Inline audio tags
Use Audio tags inside segment text for local vocal actions, alongside segment instructions for pace, emphasis and emotional progression. Chinese uses tags such as[叹气] and [停顿]; English uses (sighs) and (pause). Use the exact language-specific spellings and keep tags sparse. A pause tag does not specify an exact duration.
Video Lab sends tags unchanged to TTS and preserves supported tags when splitting long text. Exported SRT removes exact supported tags for the segment’s language while retaining ordinary bracketed text. Caption timing remains estimated within audio pieces, including time occupied by vocal actions; review synchronization before publishing. Custom HTML or burned-in captions should likewise exclude control spellings without deleting ordinary bracketed prose.
Generated-image review
When the host agent can view images, the bundled video Skills direct it to inspect every selected generated source before composition. Review visible anatomy, hands and object contact, scale, perspective, scene logic and continuity against the script and intended style. Intentional stylization, disability, occlusion and fictional settings are not automatically defects. Repair obvious unintended problems and inspect the replacement, then check the composed crops and motion snapshots. Review outcomes belong in the source ledger and storyboard. If image viewing is unavailable, the agent records that limitation and continues other work without claiming visual approval. This uses the host’s existing vision capability; the CLI does not add an image-analysis service.video inspect remains a technical check and does not certify anatomical or narrative correctness.
