Back to the blog

SubtitleGenerator editorial

Speaker label format for SRT, VTT and plain transcripts

How to format speaker labels in subtitle files and plain transcripts — the WebVTT voice tag, the SRT name prefix, JSON fields, and the consistency rules that matter more than the exact style.

SubtitleGenerator editorial Updated October 4, 2026

Once a transcript has more than one voice, every line needs an attribution — and the attribution has to be formatted somewhere. Unlike timing or text, speaker labels do not have one universal standard: plain transcripts, SRT, WebVTT and JSON each carry them differently. Here is the landscape, and the rules that actually matter.

Plain transcripts: a prefix at line start

The dominant convention in plain text and script-style transcripts is the simplest one: label, colon, space, then the line.

Interviewer: Tell me how the idea started.
Maya: We kept re-watching the same clip.

It is readable, searchable, and survives any editor. Variants add a line break after the label or bold the name in markdown — none of them change the meaning, so pick one and keep it consistent for the whole document.

SRT: a text prefix, because there is no native field

SRT's cue format carries only an index, a timecode line, and text — there is no speaker field to fill. The convention is therefore the same prefix, written into the cue's text:

3
00:00:14,000 --> 00:00:17,200
Maya: We kept re-watching the same clip.

This is what SubtitleGenerator's SRT export produces for cues with an assigned speaker; cues you have not assigned stay unprefixed, so the label only appears where you put it.

WebVTT: the native voice tag

WebVTT is the format with real support for speakers — the voice span:

3
00:00:14.000 --> 00:00:17.200
<v Maya>We kept re-watching the same clip.</v>

The tag keeps the label semantically separate from the spoken text, which players and editors can use for rendering or styling. SubtitleGenerator's VTT export uses <v Name> for assigned speakers for exactly this reason — the label is metadata, not glued-on text.

JSON: a field, not a convention

Machine-readable exports can carry the label as data. SubtitleGenerator's JSON export includes a speakerName field on each cue (and the cue text stays clean), which is the right shape for anything downstream that wants to group by speaker without parsing prefixes out of text.

The rules that matter more than the style

  • One voice, one label, everywhere. "Maya" and "Speaker 2" must not mix in the same document — rename once, globally.
  • Real names for published work. "Speaker 1" is an intermediate state; interviews and podcasts get quoted, and quotes need names (or role labels like Interviewer and Guest).
  • Attribution per cue, not per paragraph. Merged dialogue lines hide who said what, which is the exact problem labels exist to solve.
  • Uncertainty stays visible. If you are not sure who said a line, leave it unassigned or flagged rather than guessing — a wrong label is worse than an absent one.

If your labels come from automatic diarization, the labels start anonymous (Speaker 1, Speaker 2) and the correction pass — rename, merge, reassign — turns them into the format above. The formatting is the easy part; the consistent, human-verified mapping is what makes a transcript quotable.

Frequently asked questions

What is the standard speaker label format?

There is no single standard — plain transcripts usually write the label at the start of the line followed by a colon ("Maya: ..."), subtitle formats have their own conventions, and style guides differ on capitalization and whether to use real names or Speaker 1/2. What matters is consistency within one document and a format your player or editor understands.

How do I format speaker labels in an SRT file?

SRT has no native speaker field, so the convention is a text prefix: put the label at the start of the cue's text followed by a colon and a space ("Maya: I started in 2015"). SubtitleGenerator's SRT export writes assigned speaker names exactly this way, and leaves cues without an assigned speaker unprefixed.

Does VTT support speaker labels natively?

Yes — WebVTT defines voice spans: <v Maya>text</v> marks the cue's text as spoken by Maya. Players can use it for styling or rendering, and it keeps the label separate from the spoken text instead of gluing a prefix onto it. SubtitleGenerator's VTT export uses the native voice tag for assigned speakers.

Should I use real names or Speaker 1 and Speaker 2?

For anything you will quote, publish or edit — real names (or role labels like Interviewer and Guest); they make the transcript searchable and readable. Speaker 1/2 is fine as an intermediate step while you confirm the mapping, and it is the honest output of automatic diarization, which groups voices without knowing names.

Put the workflow on a video

Inspect a local file or open the built-in sample to preview caption styles and export formats. Your video stays on your device.

Upload your video