How we format your content
What you'll actually see in a delivered transcript, caption file or dub - and how to tell us to do it differently.
Every order gets an AI-drafted transcript and translation, reviewed and corrected by a real human editor before it's marked ready for delivery - never shipped un-reviewed. What follows is how that editor formats things by default. You don't have to accept our defaults: pick a different transcript style when you submit an order, or upload your own style guide or reference document and we'll follow it instead, wherever it conflicts with what's below.
1. How exact your transcript is
Pick one when you submit an order (Advanced settings on the order form):
- Clean read (the default)
- Your words, cleaned up into standard grammatical writing - filler words ("um," "you know," "sort of") and false starts that don't add anything are trimmed, but nothing you actually said is omitted or paraphrased away. This is what most subtitle and dub clients want, since it reads and sounds most natural.
- Verbatim, lightly cleaned
- Keeps natural speech patterns and most filler words, only trimming the most obviously excessive ones.
- Full verbatim
- Every false start, stammer and filler word, exactly as spoken - the standard for legal and research use, where the transcript is evidence of exactly what was said, not a readable summary of it.
2. Speaker labels
If you ask for speakers to be labeled, each one is identified by name where we can determine it (from
your file, its description, or on-screen text), or a generic label like SPEAKER 1: where
we can't. A label appears once, at the start of that speaker's line - it's never repeated on every
line, and in a dub it's never spoken aloud by the voice track, the same way a real voice actor
wouldn't say a character's name before their own line.
3. Formatting you'll see in the text
- Cut-off speech
- A trailing
--where a speaker was interrupted or cut themself off, rather than an ellipsis or a guess at how the sentence would have ended. - Inaudible speech
- Tagged
[inaudible]rather than a guessed-at word we can't actually make out. - Crosstalk / people talking over each other
- Prefixed with
>>. - Music, sound effects & non-speech audio
- Noted in brackets where it matters to following along, e.g.
[MUSIC],[APPLAUSE]- not transcribed as if it were speech. - Italics
- Only used if you ask for it (Advanced settings) - typically for on-screen text or off-screen narration. It's real formatting that carries into your delivered caption files, not just a visual note.
4. Captions & subtitles
Delivered .srt/.vtt caption files are capped at 6 seconds on screen per
caption and wrapped at roughly 32 characters per line, so nothing sits on screen too long to read or
runs past what fits comfortably in a subtitle. This is applied automatically to every delivered
caption file - not something you need to request or manage.
5. Translation & dubbing
A translation aims to say what you actually meant in natural target-language phrasing, not a word-for-word substitution that might read stiffly. A dubbed voice track is timed to the original speaker's own pacing wherever the language allows it - see our blog for more on how we handle a translation that runs noticeably longer or shorter than the original audio it needs to fit.
6. Want it done differently?
Three ways to get something other than the defaults above, in order of how specific they are:
- Pick a different transcript style and toggle speaker labels, italics, and lyric transcription right on the order form - no extra file needed.
- Upload your own style guide or reference document when you submit an order (terminology, brand names, spelling preferences, formatting rules) - we'll follow it over our own defaults wherever the two disagree. You can see it was received, and re-download it any time, from that order's page.
- Tell us directly on the message thread of any order, before or after delivery - a real person reads it, not a bot, and we can revise a delivered file if something doesn't match what you actually wanted.
This page covers the output-facing conventions only - the terminology and mechanics our own editors use internally are a separate, much longer reference we keep for staff training, not something a client needs to read to know what they're getting.
Questions before you submit a file?
A real person reads these, not a support queue bot.