Subtitling is one of the clearest cases where automation moved the work rather than removing it. The transcript is solved. The timing, the segmentation and the compression are not, and those are what a viewer experiences as good or bad subtitles.
What to put in the public brief
- Video length and language
- Whether translating or only timing
- The output format
- Number of speakers
- Whether captions for the deaf and hard of hearing are required, which includes sound events
Keep the video private.
What to put in the private instructions
Give the video, the auto-caption draft and the format
Say the file type and the frame rate. A subtitle file at the wrong frame rate drifts, and the drift only appears near the end.
Set the reading rate and maximum lines
Two lines, never three. State the characters-per-second target so shortening a line is an instruction rather than a liberty.
Say to break on sense, not on width
"Where a speaker would pause" beats any rule about line length, and it is what separates subtitles that read easily from subtitles that are merely present.
Say whether to caption sound
Music, laughter, a door. Required for accessibility captions, wrong for ordinary translation subtitles, and never guessable.
Proof worth requiring
Ask for
The subtitle file in the named format, and one screenshot of the busiest passage with subtitles displayed so you can see the worst case rather than the average.
Do not ask for
A re-exported video with subtitles burned in, unless that is the deliverable. You lose the ability to correct a single line without a full re-render.
When this task type is the wrong tool
If you need the spoken content as text rather than on screen, that is transcription and it costs less.
If the video will run in several languages, subtitle the source language first and have that file translated. Subtitling each language from the audio separately pays for the timing work over and over.