SubtitleGenerator editorial
Why we built automatic speaker labels into a subtitle tool
Interviews and podcasts fail subtitles in a specific way: two voices, one transcript. Automatic speaker labels now ship in SubtitleGenerator — with the same flagged-doubts review flow as the words.
When two people talk over each other, subtitles stop being a transcription problem and become an attribution problem. The words are usually right — who said them is what breaks. A 90-second interview clip comes back as one wall of text, and every "yeah", "right" and "absolutely" belongs to nobody.
Today that changes: SubtitleGenerator now attaches automatic speaker labels to transcriptions — on every plan, including free, with no separate mode to switch on. Upload an interview, and the editor groups the lines by speaker: a legend lists everyone it found, each cue carries that speaker's color in the list and on the timeline, and cues the model is unsure about are flagged the same way uncertain words always have been.
Built like the review flow, not like a magic button
We shipped a subtitle tool that shows you its mistakes because clean drafts hide work. Speaker labels follow the same rule:
- Doubt is displayed, not smoothed over. Cues that need a speaker check are counted in the header — "1 subtitle needs a speaker check" — and clicking it jumps you straight to the flagged cue.
- Corrections are cheap. Rename a speaker once and every cue follows. Merge two speakers into one when the model over-split a voice, with undo. Reassign a single cue when someone else actually said it.
- Human decisions win. Your confirmations and fixes persist through refreshes and restores, and a partial re-transcription does not wipe them — the editor re-applies your decisions on top of new audio evidence.
What we deliberately did not do is promise a diarization accuracy number. Diarization quality depends on the recording — clean two-person interviews cluster well; noise, overlap and long meetings are genuinely harder, and every provider has different failure shapes. Instead of a marketing percentage, the product counts what is left to review, and we measure real correction patterns to decide what to build next.
Who it is for
Anyone whose source material is a conversation: interview videos, podcasts, two-person YouTube formats, oral-history projects, research recordings. If your footage has one voice, nothing changes — the speaker surfaces simply do not appear.
Live today, on every plan
Speaker labels are part of normal transcription — upload a multi-speaker video and the editor does the rest. Free includes it (your first project has standard transcription; 60 minutes a month when you create an account), and paid plans raise the included minutes.
If you try it on a messy real-world recording, we would genuinely like to hear where it struggles — the review counter that shows what the model is unsure about works for speakers exactly the way it works for words.
Frequently asked questions
What are automatic speaker labels?
When you transcribe a video with more than one voice, the editor now groups the lines by speaker. Each subtitle row shows who is talking, the speaker legend lists everyone it found, and the timeline marks each cue with that speaker's color. It runs automatically as part of transcription — there is no separate mode to switch on.
How accurate is speaker detection?
It depends on the audio. Clean interviews with distinct voices cluster well; noisy recordings, crosstalk and long meetings are harder. That is exactly why labels ship with a review flow: uncertain cues are flagged, one subtitle needing a speaker check is counted in the header, and Rename, Merge and per-cue reassignment let you fix a whole conversation in a few clicks. We measure those corrections rather than promising a number we cannot stand behind.
Is the speaker feature free?
Yes — speaker labels are included in normal transcription on every plan, including the free tier. Your first project includes standard transcription, and paid plans raise the included minutes per month.
Does it change my video or split one track into two?
No. Speakers are labels on the single subtitle track — the timeline stays one track, cues keep their text and timing, and exported files match what you see in the editor. Speaker names appear in the editor; what gets burned into an exported video follows your subtitle style.
Put the workflow on a video
Inspect a local file or open the built-in sample to preview caption styles and export formats. Your video stays on your device.
Upload your video