What Makes a Good Subtitle? Reading Speed, Line Breaks and Timing Explained
Neither promise says anything about the thing viewers actually experience — whether the subtitle can be read. A caption track can be word-perfect and still fail, because the text arrives in blocks too long to finish, timed too fast to follow, broken in places that fight the grammar of the sentence.
The best subtitles are the ones you never notice. This guide covers the craft standards that make that possible: reading speed, line length, line breaks, timing, and the judgement calls in between. They are the same standards used by streaming platforms and broadcasters — and they apply just as much to a 40-second social clip as to a feature film.
The three properties of a subtitle that works
A good subtitle is:
- Accurate — it says what was said.
- Synchronised — it appears when the words are spoken, and leaves when they're done.
- Readable — the viewer can finish it comfortably before it disappears, without losing the picture.
Automatic speech recognition has largely solved the first property. Timing engines handle the second reasonably well. The third — readability — is where almost all automatic captioning still fails, because readability isn't a transcription problem. It's a typography-and-time problem. The rest of this guide is about solving it.
Reading speed: the number that governs everything
Reading speed is one of the most important measures of subtitle quality. It is usually expressed as characters per second (CPS), calculated by dividing the total number of characters in a subtitle, including spaces and punctuation, by the time it remains on screen. In practice, professional subtitles are generally kept around 15–17 CPS, while 20 CPS is commonly treated as an upper limit for adult viewers. Traditional European subtitling has often used a slower pace of around 12 CPS, while guidance from broadcasters such as the BBC typically corresponds to roughly 15–17 CPS. As a simple rule of thumb, multiplying CPS by 10.5 gives an approximate words-per-minute reading speed: 12 CPS is about 125 words per minute, 17 CPS about 180, and 20 CPS is a demanding pace that should be used sparingly.
The important point is that the maximum reading speed should not become the default target. Keeping subtitles consistently close to the upper limit can be tiring and forces viewers to focus more on reading than on the video itself. Aiming for around 15–17 CPS leaves room for faster sections of dialogue without making the entire experience unnecessarily demanding. Subtitle reading is also different from reading ordinary text: viewers cannot slow down, go back or skim at their own pace, because each subtitle disappears at a fixed time. When subtitles move too quickly, viewers may simply not have enough time to finish reading them, making excessive reading speed a functional accessibility problem rather than merely a matter of style.
The shape: two lines, ~42 characters, no orphans
The physical form of a subtitle is remarkably standardised:
- Maximum two lines per subtitle. Three-line subtitles cover too much picture and force vertical re-scanning.
- Maximum ~42 characters per line for Latin-script languages — Netflix's figure, now the de facto industry norm. Broadcast traditions built on teletext used 37, and many broadcasters still do.
- Keep to one line when the text fits. A second line is a necessity, not a default.
- Balance the lines. When you must break into two lines, aim for roughly even lengths. Avoid the "orphan" — one long line with a single stranded word beneath it. Where a perfectly even split isn't possible, a slightly shorter top line is generally preferred: it keeps the block compact and covers less of the image.
These limits exist for the same reason as the speed limit: the eye has to travel the text and the picture in the same seconds. Short, predictable shapes make that trip fast.
Line breaks: follow the grammar, not the ruler
Here is where good subtitling most visibly separates itself from automatic output. Where a line breaks matters as much as how long it is, because readers parse text in grammatical chunks. A break in the middle of a chunk forces the brain to hold an incomplete structure open across the jump — tiny effort, multiplied by every subtitle in the video.
The rule: never separate words that belong to the same grammatical unit.
Keep together:
- Article + noun (the / report ✗)
- Preposition + its object (to the / committee ✗)
- Verb + auxiliary or particle (has / finished, give / up ✗)
- First name + last name, numbers + units, dates
Poor break (splits the prepositional phrase):
She handed the report to
the committee chair before the vote.
Good break (each line is a complete unit):
She handed the report
to the committee chair before the vote.
Poor break (strands the verb from its object):
After the meeting, we decided to postpone
the launch until spring.
Good break:
After the meeting,
we decided to postpone the launch until spring.
Notice the good versions are not shorter — they are simply broken where the sentence itself breathes. That is the entire principle. When a sentence offers a comma, a clause boundary, or a natural pause, break there; the ruler is a constraint, not a guide.
Timing: when the subtitle lives, and the gaps between
Four numbers govern subtitle timing:
Minimum duration: ~0.8–1 second. Even a subtitle that just says "Yes." needs time to be noticed, fixated, and read. Netflix's floor is ⅚ of a second (about 833 ms); anything shorter registers as a flicker.
Maximum duration: 6–7 seconds. Past that, readers finish, re-read, and start wondering whether the player has frozen. If the speech runs longer, split the subtitle.
Sync to speech onset. The subtitle should appear as the words begin — a fraction early is tolerable, noticeably late is not, because viewers hear the voice and glance down for text that isn't there yet.
Gaps between consecutive subtitles: ~2 frames (80–120 ms). This one is invisible and essential. When one subtitle is replaced instantly by the next with no gap, the eye often fails to register that the text changed — especially if the two blocks are similar in shape. A two-frame blink tells the reader: new text, start again. Professional subtitling inserts these micro-gaps everywhere; raw auto-captions almost never do.
And one refinement that separates broadcast-grade work: respect shot changes. A subtitle that hangs across a cut tends to get re-read — the scene changed, so the brain assumes the text did too. Where timing allows, end the subtitle at the cut or carry it well past; don't let it straddle the edit by a few frames.
One thought per subtitle
Beyond the mechanics, there is segmentation: what goes in each block. The principle is that each subtitle should be a self-contained unit of meaning — a sentence, or a complete clause of one.
- Prefer whole sentences per block. When a sentence must span two blocks, split at a clause boundary, and let the first block end at a natural pause.
- Don't carry two or three orphaned words into a new subtitle. If the tail of a sentence is that short, rebalance the split.
- In dialogue, don't mix the end of one speaker's sentence and the start of another's in the same block unless you're deliberately formatting a dual-speaker subtitle (with dashes).
Bad segmentation is why some caption tracks feel inexplicably tiring even when every individual block is short: the cuts land mid-thought, so every block leaves the reader suspended.
Verbatim or edited? An honest answer
Real speech is full of false starts, fillers, and repetition: "So, um, I think — I think what we should probably do is…" Should the subtitle carry all of it?
The professional answer is: it depends on the job of the video, and there is a genuine trade-off either way.
The case for light editing. When speech outruns reading speed, something has to give, and trimming um, you know, false starts and immediate self-repetitions is the standard first move in professional subtitling. It converts unreadable 24 CPS speech into readable 16 CPS text while preserving every word that carries meaning. For marketing content, courses, and social clips, this is almost always right.
The case for verbatim. Many deaf and hard-of-hearing viewers explicitly prefer verbatim captions — the hesitations, the fillers, the way someone talks is part of what everyone else gets through their ears, and editing it out is a small act of gatekeeping. For interviews where the speaker's manner matters, for legal and archival records, and as an accessibility default when reading speed permits, verbatim is the more respectful choice.
The dishonest position is pretending there's no tension. The honest workflow is: verbatim where reading speed allows; condense only when the alternative is a subtitle nobody can finish; and never "improve" what someone actually said.
For deaf and hard-of-hearing viewers: SDH extras
Standard subtitles assume the viewer can hear everything except the language. SDH (Subtitles for the Deaf and Hard of Hearing) assumes they can't hear at all, and adds:
- Speaker identification when it isn't visually obvious — an off-screen voice, a phone call, overlapping speakers: DAVID: We need to talk.
- Sound cues that matter to the story: [door slams], [muffled argument next door], [phone buzzes]. The test is narrative relevance — caption the sound the plot needs, not every ambient noise.
- Music notation: ♪ for lyrics being sung, or a description — [tense strings build] — when the score is doing storytelling work.
If accessibility is part of why you caption — and legally, it increasingly is — SDH conventions are the bar, not the bonus. (For the legal side, see our guide to the European Accessibility Act and what it requires for video.)
Why raw auto-captions fail these standards
None of the above is exotic knowledge — so why does most automatic captioning ignore it?
Because speech recognition engines segment by pause detection, not by grammar or reading speed. The engine hears a stretch of speech, transcribes it, and stamps the whole stretch as one block. A confident speaker who doesn't pause produces a five-second, 120-character wall. The block boundaries land wherever the speaker breathed — mid-clause, mid-name, mid-number. Consecutive blocks butt against each other with zero-frame gaps. Every individual word can be correct while the track as a whole is, functionally, unreadable.
That is the precise meaning of "accurate and unusable" — and it's why a caption track's word-accuracy percentage tells you very little about whether it works.
The spec at a glance
Professional subtitles should typically stay within a reading speed of 15–17 characters per second (CPS), roughly equivalent to 160–180 words per minute, with 20 CPS treated as an absolute ceiling. Each subtitle should use one line where possible and no more than two, with a maximum of around 42 characters per line for Latin scripts. Line breaks should follow natural grammatical boundaries, while each subtitle should generally remain on screen for at least 0.8–1 second and no longer than 6–7 seconds. Consecutive subtitles should be separated by a small gap of roughly two frames, or 80–120 milliseconds, and captions should avoid straddling shot changes by only a few frames. Segmentation should follow complete sentences or meaningful clauses rather than leaving short, orphaned fragments, while subtitles for deaf and hard-of-hearing audiences (SDH) should also include speaker identification and sound cues that are relevant to understanding the narrative.
Where Inwista fits
Everything above can be done by hand. Professional subtitlers have done it by hand for decades — at real cost in hours.
Inwista automates it. Transcription gives you the accurate text; the editor gives you full manual control; and Enhance applies this article in one click: it splits over-long blocks into readable ones, breaks lines at grammatical boundaries, inserts the industry-standard micro-gaps between subtitles, and brings reading speed down toward the professional range — while you keep the final say on every block. Presets tune the output for broadcast-style or social-style delivery.
Inwista's free plan lets you run the whole workflow on your own material — upload, transcribe, structure, edit and export. Current limits and plan details are on the pricing page.
Frequently asked questions
What is CPS in subtitling? Characters per second — the number of characters in a subtitle (including spaces and punctuation) divided by its on-screen duration. It's the standard measure of subtitle reading speed. Professional work targets 15–17 CPS; Netflix allows up to 20 CPS for adult English content.
How many characters should a subtitle line have? At most about 42 characters per line for Latin-script languages, with a maximum of two lines per subtitle. Broadcast traditions rooted in teletext often use 37.
What is the six-second rule? A classic European guideline: two full subtitle lines should stay on screen for about six seconds — which works out to roughly 12 CPS. Modern platform standards allow faster speeds, but the rule survives as a reminder that comfortable is slower than possible.
Should subtitles be verbatim or edited? Both are legitimate. Verbatim preserves exactly what was said and is preferred by many deaf and hard-of-hearing viewers; light editing (removing fillers and false starts) is standard when speech is too fast to read otherwise. Condense only when reading speed forces it, and never alter meaning.
Why do auto-generated captions feel hard to read even when the words are right? Because ASR engines segment by pauses, not grammar or reading speed — producing over-long blocks, mid-clause breaks, and no gaps between subtitles. Word accuracy and readability are separate problems.
What's the minimum time a subtitle should stay on screen? Around 0.8–1 second, even for a single word. Netflix's floor is ⅚ of a second. Shorter than that reads as a flicker rather than text.
Standards cited are the published or widely documented guidelines of the named platforms and broadcasters as of this writing; individual style guides evolve, and specific productions may impose their own. Last updated: July 2026