Skip to content

Transcripts

A transcript gives Mixture the words behind a recording — each word with the exact moment it was spoken. Once a clip has one, you can edit the footage by the words instead of the waveform, and text layers can repeat over the words to become captions that re-time themselves with every cut.

Mixture doesn’t transcribe audio itself yet — honestly. The words arrive one of two ways:

  • Ask an agent. Connect a coding agent and say “transcribe the narration” — the agent transcribes the audio and attaches the words to the recording. This is the usual route.
  • Import a file. Drag a transcript JSON into the project’s assets. The format is plain word timings — subtitle files (SRT/VTT) aren’t understood yet.
The transcript format
{
"words": [
{ "word": "Welcome", "start": 0.32, "end": 0.71 },
{ "word": "to", "start": 0.71, "end": 0.84 },
{ "word": "the", "start": 0.84, "end": 0.95 },
{ "word": "ballpark.", "start": 0.95, "end": 1.52 }
]
}

Times are seconds into the source recording. Keep the spoken punctuation in the words — it’s what makes sentences readable in the editor.

A transcript belongs to the recording, not to any one clip — every clip that uses that recording, in any scene, shows the same words automatically.

Import a words file and it lands as a transcript asset in the assets view, with its word count. That matters because one transcript can serve several recordings at once: the screen capture, the camera, and the mic from one take were all rolling through the same words.

Attach it in either of two places, both doing the same thing:

  • From a clip. Select it, add the Transcript row in the properties panel, and pick a source under Use words from.
  • From the asset. Open a video or audio asset’s detail view and use Link transcript in its Transcript row.

Either way the words land on the recording, so every clip using it — now and later, including ones you split or duplicate — shows them, and so do any linked clips. The link is live: refine the transcript asset and every recording linked to it picks up the change. Unlink removes the connection; the transcript asset stays.

You can borrow another recording’s words, not just a transcript file. The list under Use words from is everything in the project that actually has a transcript — files you imported and recordings that were transcribed. Pick the mic wav and the camera reads its words. That’s the usual shape of a take: one file gets transcribed, the rest of the take points at it.

The clip’s Transcript row also tells you where its words are coming from — on this recording, or from a linked clip. If it says No words yet, picking a source is the fix.

Double-click a transcribed clip and the clip editor shows a word lane beneath the waveform — every word sits exactly under its moment of audio. Click a word to jump there. Drag a cut across a rambling sentence and its words turn struck-through red: the sentence is out of the video, and dragging the cut away brings it back. Zoomed out, the lane shows whole sentences instead of words, so it stays readable.

Clips in a linked take share words too: a camera clip linked to the transcribed mic track shows the mic’s words in its own word lane — one transcription serves the whole take.

The words are data, and the data system can use them. Set a text layer’s Repeat source to Transcript words and the layer repeats over the clip’s spoken words — one copy per word, each already timed. Cut a sentence and its words vanish from the captions; speed a stretch up and its words read faster. You never redo caption timing after an edit.

Captions from a transcript in the data guide covers the patterns — karaoke words, build-up lines, two-font emphasis.

Attaching new words replaces the old ones — so if the timings drift against the waveform, re-transcribe rather than nudging every cut to fit. Unlink in the asset’s detail view removes an attached transcript or breaks the link to a transcript asset. A clip’s Transcript override clears back to Auto at any time.