Skip to content

Captions & transcripts

Anything with spoken word — talking heads, voiceovers, podcast clips — agents handle end to end: transcript, captions, and the edits between.

Caption this clip — clean blocks, bottom third, phrase by phrase.

Karaoke captions: highlight each word as I say it, in brand colors.

Cut the ums and the dead air, and tighten the pauses.

  • Transcribe on the fly. No transcript? The agent runs speech-to-text itself and attaches word-level timings to the clip — from then on, transcript features work as if it was transcribed in-app.
  • Captions built the data-driven way. One text layer repeated over the transcript’s words or phrases, with content and timing bound per row — so a recut or a re-transcribe retimes every caption automatically. There’s no pile of hand-placed text layers to maintain.
  • The full caption wardrobe — subtitle blocks, kinetic word reveals, phrase-by-phrase pacing, per-word emphasis in a second font, lower-third speaker labels.
  • Edit by the transcript. Because words carry timings, “cut the part where I repeat myself” is a precise instruction, not a guess — the same word-level editing you have in the clip editor.

Captions are ordinary text layers and the transcript is editable — fix a word in the transcript panel and the caption updates with it.