There are 180 fonts in the editor. Almost none of the difference between a readable caption and an unreadable one comes from which of them you pick.
A caption is on screen for maybe a second and a half, at a size that varies with the viewer's phone, over footage you don't control, being read by someone who is not concentrating. That is a genuinely hostile set of conditions, and the things that survive them are not the things type specimens are designed to show off.
Contrast beats everything
The single biggest factor is whether the text separates from what's behind it, and video is the worst possible background because it changes every frame. A caption that is perfectly legible over a dark jacket disappears when the subject turns and a window fills the shot.
Which is why every default here carries something behind or around the letters rather than relying on the text colour alone:
- An outline — a dark stroke around light text. Cheap, works on anything, slightly ugly at large sizes.
- A shadow — softer, better on faces, less reliable over high-frequency detail like foliage or a crowd.
- A background plate — a solid or translucent box. Ugliest and by far the most robust. If the footage is genuinely unpredictable, use it.
Pick one. Text with no treatment at all looks best in the editor and worst in the wild, because in the editor you're looking at one frame.
Counters, not personalities
Where the typeface does matter, it's in the details you'd never notice at paragraph size.
Counters — the enclosed spaces inside a, e, o, g — are the first thing to close up when text is small, condensed, or heavily outlined. A face with open counters stays readable when the same size in a tighter face turns into rectangles.
Stroke weight matters in both directions. Too light and an outline eats it. Too heavy and the counters fill in. Semi-bold to bold is the useful band for captions; hairline weights that look elegant in a title are unusable here.
Ambiguous pairs are worth checking once: capital I, lowercase l, and the digit 1. Some otherwise excellent geometric faces render all three as an identical vertical stroke, which is fine in prose where context disambiguates and not fine in a caption reading "Item l of 1".
Choose a face for how it behaves at 1.5 seconds and 40 pixels, not for how it looks in a specimen at 200.
Line length and the three-word rule
The default here is three words a line, and that's a readability decision rather than an aesthetic one.
Reading is saccadic — the eye jumps, fixes, jumps. A short line is taken in with one or two fixations, which is about all you get before the caption changes. A long line forces a horizontal scan, and mid-scan the text is replaced. That's the feeling of captions being "too fast" when in fact they're correctly timed and simply too wide.
Two to four words is the band. Below two it feels stroboscopic; above four you lose people whose reading speed is anywhere below average — which includes anyone tired, distracted, reading a second language, or dyslexic.
What about a dyslexia font?
The specialist typefaces marketed for dyslexia have not held up well in controlled studies; the measurable gains generally come from what they do besides the letterforms — more spacing, shorter lines, more contrast.
You can have all three without changing typeface, and they help every reader rather than one group. That's the better trade, and it's what the defaults do.