Automatic speaker labels

Automatic speaker labels for every voice in your video

When several people talk in one video, every voice gets a speaker label in your transcription, automatically. Cues the model is not sure about are flagged for your review — you confirm, rename, merge or reassign — and the result stays in the same browser editor you already use. Free to start; your video never leaves your device.

How automatic speaker labels work

Transcription and speaker labeling run in one pass. Each spoken cue is grouped by voice, so a two-person interview reads as a dialogue instead of a wall of text. You see who is on screen in three places at once: a legend listing every speaker, a color dot on each subtitle, and a marker on the timeline.

Review what the model is not sure about

Automatic does not mean unreviewed. Cues where voices are similar or the split is uncertain are flagged, and the editor counts them for you — “N subtitles need a speaker check” stays visible until the queue is clear. Each flagged cue carries a Confirm button; naming a speaker and confirming takes one click per cue. The labels are yours: rename speakers to real names, merge two labels into one, or reassign any single subtitle.

Make each speaker visible

Give every speaker a text color and an accent color, with a contrast warning if a choice gets hard to read. Turn on “Show speaker names on video” to see the name beside each line while you proof the dialogue. One voice detected? Add a speaker manually and assign lines yourself — the review flow works the same way.

Your video stays yours

Speaker labeling runs on the same local-first pipeline as the rest of the editor: your video stays on your device, and only the extracted audio needed for transcription is processed. The labels, colors and review decisions live in your draft in this browser — nothing to install, no account required to start, and no copy of your footage on anyone else's server. Switch browsers or devices and you start a fresh project — the draft, like the video, stays on the machine it began on.

Where speaker labels earn their keep

Any recording with more than one voice benefits: a two-person interview filmed for an article, a video podcast or panel clip, a webinar, a lecture capture, a meeting recording you need to quote accurately. They all share the same shape — several voices, one recording, and a reader who has to know who said what. Automatic speaker labels turn that recording into a quotable dialogue, and the flagged-cue review turns the cleanup into a bounded pass instead of an open-ended rewrite.

Change your mind without starting over

Renaming a speaker updates the name everywhere it appears. Merging two speakers keeps every cue and folds the histories together. Reassigning moves one subtitle at a time, so a single wrong line never forces a redo. Every speaker edit shares the same undo history as your subtitle edits, and the speaker state persists with your draft in this browser — close the tab, come back, and the review queue is where you left it.

Speaker label in transcription exports

Subtitle exports (SRT and VTT) carry the corrected text and timing for every cue. Speaker names are a preview layer: they guide your review in the editor and on the video preview, and exported files match the text and timing you approved. Burned-in video exports follow the export rules of the shared editor.

Questions about speaker labels

What is a speaker label in a transcription?

A tag that says who is speaking each line — “Speaker 1”, a real name, any label you give. In a two-person interview it turns a flat block of text into a readable dialogue.

Are the speaker labels always right?

No, and the editor does not pretend they are. Uncertain cues are flagged and counted, and you confirm each one. You can rename, merge and reassign at any time.

Is the speaker labeling free?

Your first project is free, including the standard transcription with automatic speaker labels. Longer videos and more exports move to Pro, Max or pay-as-you-go.

Do noisy or overlapping voices work?

Distinct voices label best. Heavy background noise, crosstalk and long meetings with similar voices are harder and produce more flagged cues — the review flow, not a promise of perfection, is the product answer.

Can I edit a speaker label after I close the editor?

Yes. Reopen your draft in the same browser and the labels come back with it — rename, merge, reassign or confirm anything you left flagged, then export again.

Does speaker labeling change my video?

No. Your footage is untouched. The labels live in the subtitle layer: they can show beside each line in the editor preview, and exports follow the subtitle export rules.

Read why we built automatic speaker labels into a subtitle tool, start from auto subtitles, or pick a look in caption styles and the subtitle design guide before you export.

Start your speaker-labeled transcript

Add a video, let the voices separate themselves, and settle the flagged cues in one review pass — in your browser, free to start.

Upload your video