Skip to content

Caption a talking head

Captions that actually follow the audio — word-timed, styled, and built the data-driven way so a recut doesn’t strand them.

  • A clip with speech — talking head, voiceover, podcast excerpt.
  • A connected agent.

Caption this clip. Clean subtitle blocks, bottom third, phrase by phrase.

Or with more personality:

Give me karaoke captions — highlight each word as I say it, brand colors.

If the clip has no transcript yet, the agent makes one on the spot — running speech-to-text itself (a local Whisper, or an API it has access to) and attaching the word-level timings to the asset, exactly as if it had been transcribed in-app. Then it builds the captions the data-driven way: one text layer repeated over the transcript’s words or phrases, with content and timing bound per row.

That one-layer structure is the point: re-transcribe, trim, or cut the clip and the captions retime themselves — there’s no pile of hand-placed text layers to fix.

The caption layer is normal text — restyle it in the properties panel, and edit any wording via the transcript so the text and timing stay in sync. Ask the agent for treatments (“make the emphasised words pop in a second font”) and it extends the same repeat rather than starting over.