All posts
ACCESSIBILITY7 min read

Auto-captions are not accessibility

Automatic transcription gets you to about 95% word accuracy. Accessibility guidance asks for near-perfect. Everything interesting lives in that last five percent.

FathiCo-founder
Testing default caption placement against the safe areas of four different apps.

A customer asked us last month whether switching on automatic captions made their videos compliant. The honest answer is no, and the reason is worth more than a yes would have been.

Transcription runs at roughly 95% word accuracy on clean audio — one speaker, a decent microphone, no music bed. That sounds excellent until you work out what it means. A three-minute video is about 450 words. Five percent of that is around twenty wrong words, and they don't distribute themselves politely across the transcript. They cluster exactly where the audio gets hard.

The accuracy ceiling

Speech models are trained on what people usually say, which is precisely why they fail on what your video specifically says. A brand name becomes two ordinary words. A price of $4.99 becomes "four ninety nine". Someone called Siobhán becomes almost anything.

The industry benchmark most captioning guidelines point at is 99% — near verbatim, because below that the errors stop being cosmetic. The gap between 95 and 99 isn't a rounding error. It's the difference between a transcript that helps and one that misleads. A deaf viewer reading "we do not recommend this dosage" as "we now recommend this dosage" is worse off than one reading nothing.

Captions that are 95% right are not 95% accessible. They're a document somebody has to fact-check while they watch.

What breaks first

The errors are not random, and after enough of them you can predict where they'll be. Four categories cover most of what we see people correcting:

  • Proper nouns. People, brands, place names — anything the model has no reason to expect.
  • Numbers and units. Prices, dates, dosages, model numbers. Often transcribed as words when they need to be digits, or the reverse.
  • Overlapping speech. Two voices inside the same fraction of a second. Models resolve this badly and confidently.
  • Domain vocabulary. Anything rarely written down in the training data.

None of that is fixable by a better model in the general case, because the information the model is missing is yours.

Show people where to look

Reading a transcript takes about as long as the video, and that is the whole cost of getting to near-perfect. Three and a half minutes of attention on a three-minute video, because you already know the names, the numbers and what you actually said.

So we stopped hiding the transcript. It sits beside the video, the word being spoken is underlined as it plays, and the transcript is editable in place.

More usefully, the model tells us how confident it was about each word, and we keep that. Anything below 0.6 confidence gets an amber underline in the transcript panel. It isn't clever — it's a way of turning "read all 450 words carefully" into "check these eleven", which is a job someone will actually do.

What we changed

We don't describe this product as WCAG compliant, and we won't. What we ship is a very good first draft plus the fastest tools we could build for the pass that makes it true: word-level timing, the confidence underline, and a transcript you can edit without leaving the editor.

Automatic captions are an excellent starting point. Sold as anything more, they're a promise somebody else has to keep.

Fathi · Co-founder

Builds Add Captions. Writes here about the parts of captioning that turn out to be harder than they look.

More writing

All posts

Captions, timed to the word

Everything above is how Add Captions works. Upload a video and see it — five dollars a month, and you can caption as many as you like before paying anything.

Add captions free