Back to the blog

SubtitleGenerator editorial

How to correct speaker labels in a transcript

Wrong speaker names in a transcript are cheap to fix — rename once, merge over-split voices, reassign single lines, and confirm flagged cues. Works on your local video, free to start.

SubtitleGenerator editorial Updated October 4, 2026

A transcript can be word-perfect and still be wrong. Two people talk, the text comes back clean, and every line is attributed to nobody — or to "Speaker 1" and "Speaker 2" when you need real names. If you are going to quote the recording, edit a video, or publish the transcript, the labels have to be corrected by a human, because the model can only group voices; it cannot know that the second voice is called Maya.

Here is how to correct speaker labels in a transcript — first the four things that are ever wrong with them, then the exact fixes, then how the corrections behave in exports.

When speaker labels need correcting

Almost every fix falls into one of four cases:

  • Anonymous labels. The transcript says Speaker 1 and Speaker 2. Technically correct, useless for quoting. You want real names.
  • One voice, two labels. The model split a single speaker in two — often after a pause, a microphone change, or a stretch where the voice sounds different. Every line of that person is scattered across two names.
  • Two voices, one label. The opposite: a short interjection ("right", "exactly") got folded into the other speaker's label, so one name owns lines that belong to two people.
  • A single wrong line. The grouping was right all along except for one cue — someone talked over someone else, or a very short backchannel got attached to the wrong voice.

Automatic labeling surfaces the risky ones instead of hiding them: cues the model is unsure about are flagged and counted, so you know where the corrections are needed before you read every line.

How to correct speaker labels in SubtitleGenerator

All four fixes live in the same review flow as the flagged words, and they share one undo history with your subtitle edits.

Rename a speaker once

Open the speaker legend, choose the label, and rename it — "Speaker 2" becomes "Maya" everywhere it appears, in the legend, on the timeline markers and beside each line on the video preview. Renaming is the fix for anonymous labels and takes one edit no matter how many cues carry the name.

Merge an over-split voice

When one person is scattered across two labels, merge them: pick the label to fold away, confirm, and every cue it owned follows the surviving label. The editor tells you how many subtitles will move before you commit, and the merge is undoable.

Reassign the single wrong line

For a cue that belongs to a different speaker, open the cue in the inspector and pick the right speaker from the list — or clear it to unassigned if you are not sure yet. One line moves; nothing else changes. This is the fix for crosstalk and short interjections.

The under-split case has one extra step: if the second voice does not exist as a label yet, add a speaker first ("Add speaker" in the legend), then reassign the folded lines to the new label. Adding a speaker is also the manual fallback when the video has more voices than the model found.

Confirm the flagged cues

Cues where voices are similar or the split is uncertain are flagged, and the header counts them — "3 subtitles need a speaker check". Review each one and press Confirm when the assignment is right. The count drops to zero only when every flagged cue has been settled, so "all clear" means something.

Every speaker edit persists with your draft in the browser: close the tab, come back, and the labels and the review queue are where you left them.

Correcting speaker labels in any transcript

The same order of operations works whatever tool produced the transcript. Fix the systematic problems first — rename anonymous labels, merge over-split voices — because they touch many lines with one decision. Then sweep the single-line errors, because each reassignment is a judgment call. Finally, proof the labels against the audio at the points where speakers change, which is where attribution errors cluster. If your tool exports a plain text transcript, do the correction there before formatting; if it exports subtitles, correct the labels in the editor where the cues are timed against the video.

Do the corrections carry into exports

Yes. In SubtitleGenerator, subtitle file exports carry each assigned speaker's name: SRT, SBV and TXT prefix the line — Maya: I started in 2015 — WebVTT uses its native voice tag (<v Maya>), and the JSON export includes a speakerName field. Cues you have not assigned stay unprefixed, and no correction moves a cue's timing. Burned-in video exports follow your subtitle style, as they always have.

A transcript with corrected labels is quotable, searchable and accessible — and with the review flow doing the bookkeeping, the correction pass is a bounded sweep, not a rewrite. Automatic speaker labels come with every transcription, free to start, and the launch notes explain the design decisions behind them.

Frequently asked questions

What does "correct speaker label" mean?

It means fixing who is attributed with each line in a transcript. Automatic speaker detection groups lines by voice, but it does not know real names, and it can over-split one voice into two labels or give a single line to the wrong speaker. Correcting labels means renaming speakers to real names, merging labels that should be one person, reassigning individual lines, and confirming the cues a model flagged as uncertain.

How do I fix a wrong speaker name in a transcript?

Rename the label once and every line that carries it updates. In SubtitleGenerator's editor, open the speaker legend, choose Rename, and type the real name. For lines that belong to a different speaker entirely, reassign that single cue from the inspector — no need to redo the whole transcript.

Does correcting speaker labels change the transcript timing?

No. Speaker labels ride on top of the subtitle track — renaming, merging and reassigning never moves a cue or changes its text. Every speaker edit shares the same undo history as your subtitle edits, so you can step back if a fix goes the wrong way.

Do speaker label corrections carry into the exported file?

Yes. Subtitle file exports carry each assigned speaker's name: SRT, SBV and TXT prefix the line with the name, WebVTT uses its native voice tag, and the JSON export includes a speakerName field. Cues you have not assigned stay unprefixed, and timing is unchanged.

Put the workflow on a video

Inspect a local file or open the built-in sample to preview caption styles and export formats. Your video stays on your device.

Upload your video