Screen recording accessibility: captions, transcripts, contrast and pace
Most of what makes a screen recording accessible happens before you edit anything: saying what you click, keeping text readable, and not rushing. Captions and a transcript finish the job, and WCAG asks for less extra work than most people fear.
By the VeoRec team · · 12 min read
In short
An accessible screen recording has accurate captions, says out loud what is shown on screen so people who cannot see it can follow, keeps text and overlays large and high in contrast, moves at a pace people can follow, and never flashes. Publish a transcript with it, ideally a descriptive one that includes visual details. Under WCAG, captions for prerecorded video are Level A, and if your narration covers all the important visual information you usually do not need a separate audio description.
- Narrate what you do by name and location ("Export, top right") instead of "click here"; it is the cheapest accessibility fix there is.
- Captions are required for prerecorded video at WCAG Level A; check automatic captions for names, numbers and jargon.
- Publish a transcript next to the video; a descriptive transcript also covers people who are both deaf and blind.
- Zoom the screen so text stays readable, and give overlays at least 4.5:1 contrast with what is behind them.
- Slow the cursor, pause on new screens, and never let anything flash more than three times a second.
- Keep music out from under narration, or at least 20 dB quieter than your voice.
Screen recording accessibility sounds like a compliance project, but most of it is just recording well. A video where you say "click Export in the top right" instead of "click here" works for a blind customer, a customer with the sound off reading captions, and a customer who is simply looking at a different part of the screen. A video zoomed to a readable size works for someone with low vision and for everyone watching on a laptop in a small player.
The practical side starts with knowing who the recording has to work for and what WCAG actually asks of prerecorded video. After that it is mostly habits: narrating so you rarely need separate audio description, checking captions, publishing a transcript, keeping the screen readable and the pace humane. There is a checklist at the end to run before you send anything.
Who an accessible screen recording is for
It helps to be specific, because each group needs something different from the same video. The W3C's guide to making audio and video accessible frames it the same way: captions for people who are deaf or hard of hearing, description of visual information for people who are blind or have low vision, transcripts for both and for people who are deaf and blind (W3C WAI, Making Audio and Video Media Accessible).
| Who | What gets in their way | What helps |
|---|---|---|
| Deaf or hard of hearing | Narration with no captions; important sounds (an error beep) not mentioned | Accurate captions; transcript; on-screen text for key values |
| Blind | Narration that says "this" and "here"; silent stretches where things happen on screen | Narration that names every action and result; descriptive transcript |
| Low vision | Small text, low contrast, a thin cursor, fast scrolling | Zoomed screen, high-contrast overlays, highlighted clicks, slower movement |
| Deaf and blind | Any content only available as sound or picture | A descriptive transcript they can read on a braille display |
| Cognitive or learning disabilities | Fast speech, jargon, long videos with many topics | Plain words, one task per video, pauses between steps, chapters |
| Photosensitive epilepsy | Flashing content | Nothing that flashes more than three times a second |
| Everyone else, sometimes | Open office without headphones, a noisy train, a second language | All of the above, especially captions and a transcript |
That last row is why this is not a niche concern for a support or product team. The customer reading captions on a train and the customer who is deaf are served by the same captions.
What WCAG asks of a recorded video
The Web Content Accessibility Guidelines are the standard most accessibility laws and procurement rules point to. Internal team videos may not be formally in scope where you work; customer-facing help videos, product tutorials and anything on your public website usually are. These are the success criteria that matter for screen recordings, from the WCAG 2.2 Understanding documents.
| Criterion | Level | What it means for a screen recording |
|---|---|---|
| 1.2.1 Audio-only and Video-only (Prerecorded) | A | A silent recording (no narration) needs a text alternative or an audio track describing what happens. |
| 1.2.2 Captions (Prerecorded) | A | A narrated recording needs captions for the speech and important sounds. |
| 1.2.3 Audio Description or Media Alternative (Prerecorded) | A | Visual information must be available another way: audio description or a full text alternative. Not needed if the narration already conveys it. |
| 1.2.5 Audio Description (Prerecorded) | AA | Audio description is required at AA, again unless the soundtrack already conveys all the important visual information. |
| 1.4.2 Audio Control | A | If audio plays automatically for more than 3 seconds, people must be able to pause it or control its volume. A reason not to autoplay embeds. |
| 1.4.3 Contrast (Minimum) | AA | Text, including text in video such as overlays, needs 4.5:1 contrast (3:1 for large text). |
| 2.3.1 Three Flashes or Below Threshold | A | Nothing flashes more than three times in any one second. |
Each criterion has a W3C Understanding document that explains its intent and exceptions: 1.2.1, 1.2.2, 1.2.3, 1.2.5, 1.4.2, 1.4.3, 2.3.1.
Read the exceptions in 1.2.3 and 1.2.5 again, because they are good news for anyone making screen recordings. The Understanding document for 1.2.3 says that if all the important information in the video is already conveyed in the audio track, no additional audio description is necessary. For a narrated screen recording, that is within your control. The next section is about how to do it.
Narrate what is on screen so nobody needs to see it
W3C calls this integrated description: the speaker describes the relevant visual information as part of the main narration, so no separate description track is needed. Their guidance notes that it works well for presentations and instructional videos, and costs nothing extra if you plan it from the start (W3C WAI, Description of Visual Information). Screen recordings are the ideal case.
In practice it comes down to a few habits:
- Name the control and where it is. "Click Export, in the top right corner of the table", not "click this".
- Say the result. "A dialog opens asking which columns to include." Without that, a listener does not know anything happened.
- Read out values that matter. "The total now shows 1,240 dollars", not "and now the total is correct, as you can see".
- Describe errors. "A red message under the field says the date format is invalid." Error states are exactly what people need to recognize.
- Do not rely on color alone. "The green one" means nothing to someone who cannot see green. "The Active badge" works for everyone.
- Mention important sounds if your app makes them, such as a notification chime that means the upload finished.
Compare the same step, from a support recording about exporting invoices, narrated both ways.
The second version is longer by a few seconds, and it is better for every viewer, not only blind ones. It also makes captions and the transcript useful on their own, since the transcript of the first version is meaningless. W3C's content guidance gives the same advice with a different example: describe the thing itself rather than saying "like this" (W3C WAI, Audio Content and Video Content).
Plan the description into your outline
Integrated description is much easier to plan than to improvise. W3C's description guidance says it plainly: when writing the script, make sure all relevant visual information is included (W3C WAI, Description of Visual Information), and its planning guide is a good way to work out which other features a given video needs. For a screen recording, that means adding the location and the result to each bullet of your outline: not just "Download, PDF" but "Download, top right next to Print; menu opens; choose PDF; message at bottom". You will say it naturally because it is already in front of you.
Captions: switch them on, then check them
Captions are a text version of the speech and of the non-speech sounds needed to understand the video. For screen recordings that usually means your narration plus any meaningful sound from the app. The W3C has a short video on who depends on captions and what makes them work; it is captioned, as you would hope.
Web Accessibility Perspectives: Video Captions (video, W3C Web Accessibility Initiative (WAI))
Automatic captions have made this far easier, but they are not finished captions. The W3C's caption guidance says automatically generated captions do not meet user needs or accessibility requirements unless they are confirmed to be fully accurate, and gives an example where "4 to 5 minutes" was transcribed as "45 minutes" (W3C WAI, Captions/Subtitles). In product recordings, the words automatic captions get wrong are predictable: product and feature names, people's names, acronyms, numbers and anything you said quickly.
- Watch once with captions on and the sound off. If you cannot follow, neither can a deaf viewer.
- Fix product names, numbers and technical terms first; those carry the meaning.
- If you cannot correct a caption in your tool, correct the transcript or the written steps published with the video, and re-record a sentence if a wrong number would mislead.
- Speak clearly and avoid talking over yourself; clean audio produces better automatic captions in the first place.
Every VeoRec recording gets automatic captions and a transcript, on the Free plan too, and captions start switched on for viewers, so nobody has to find a CC button before the first sentence. Pro can translate captions and transcripts for viewers in other languages. Treat all of it as the first draft: still worth a check before a recording goes to customers.
Publish a transcript, and make it descriptive
A transcript is the whole video as text: speech, important sounds and, in a descriptive transcript, the visual information too. W3C recommends descriptive transcripts because they serve the widest range of people, including those who are both deaf and blind, who may read them with a braille display (W3C WAI, Transcripts).

If you narrated well, your transcript is already most of the way to descriptive. To finish it:
- Break it into paragraphs at each step, and add headings for longer recordings.
- Add anything visual the narration skipped: a value on screen you did not read out, a label, an error message.
- Remove filler words ("um", "so, yeah") that make reading slow without adding meaning.
- Put the transcript, or a clearly labeled link to it, directly under the video, not on a separate page nobody finds.
Transcripts have a side benefit: they make the video searchable, both in your help center and across your recording library. For customer-facing videos, knowledge base videos covers how to place transcripts and written steps next to an embed.
Make the screen readable at small sizes
Your monitor is bigger and sharper than most viewers' players. A recording of a dense dashboard at 100 percent zoom on a large screen becomes unreadable in a 640 pixel embed, and for someone with low vision it is unreadable at any size.
- Zoom the browser to 125 percent or more for dense interfaces, and record a single tab or window rather than a large full screen.
- Record at 1080p so small text stays sharp when viewers zoom or go full screen.
- Make clicks visible. Click highlights and a larger cursor help everyone follow, and they matter most for people with low vision. VeoRec can highlight clicks while you record.
- Use drawings to point, not to decorate. Circling the field you are talking about is useful; scribbling across the screen is not.
- Check overlay contrast. Text you add in editing needs at least 4.5:1 contrast with what is behind it under WCAG 1.4.3 (3:1 for large text). White text on a pale screenshot fails; put it on a solid dark band.
- Prefer the theme with the clearest text. Some apps' dark modes use gray-on-gray text that is lower contrast than their light mode. Check before you record.
A camera bubble can also get in the way: if it covers a menu or a button you click, viewers lose the very thing they need to see. Keep it small and in a corner that is not part of the task, or turn it off for pure how-to recordings.
Pace, motion and flashing
Speed is an accessibility issue. People with cognitive or learning disabilities, people working in a second language and people using captions all need time to connect what they hear or read with what they see.
- Pause on every new screen for a second before you talk about it.
- Move the cursor in straight lines and stop on the target before clicking. Circling and wiggling make it harder to follow.
- Scroll slowly, or better, scroll and then stop before you talk about what is on the page.
- Keep each video to one task. Shorter videos are easier to follow and easier to rewatch. If it has several parts, add chapters.
- Use plain words and explain any acronym the first time. W3C's content guidance makes the same point about jargon and idioms.
- Never flash. Rapid toggling, flickering loaders or quickly cut sequences can trip WCAG 2.3.1. If something on screen flashes, cut it or blur it.
Chapters are an underrated accessibility feature. They let people jump back to the step they missed instead of scrubbing. In VeoRec you can drop a chapter marker with the M key as you reach each step and name it on the finish screen.
Keep the audio clear
For a screen recording, the narration is the content. Anything that competes with it hurts people who are hard of hearing first, and everyone else soon after.
- Skip background music, or keep it at least 20 decibels below speech, as W3C recommends for media where speech is the main content.
- Record in a quiet room with the microphone close to your mouth. Hum, echo and keyboard clatter make speech harder to understand and captions less accurate.
- One voice at a time. If two people talk in a recording, identify who is speaking in the transcript.
- Do not autoplay embedded videos with sound. WCAG 1.4.2 requires a way to pause or control audio that plays automatically for more than three seconds, and autoplay talks over screen readers.
Check the link and the player you share
All of this work is wasted if the viewer cannot operate the player. W3C's page on media players asks for keyboard support, a visible focus indicator, clear labels and enough contrast on the controls, and describes extras that help many viewers, such as changing the playback speed, adjusting how captions are displayed and interactive transcripts (W3C WAI, Media Players). Before you standardize on a recording or hosting tool, try its player with only a keyboard and with your operating system's screen reader.
Also check what the link asks of the viewer. A sign-in wall in front of a help video is an extra barrier, and every extra form is one more place where a keyboard or screen reader user can get stuck. With VeoRec, viewers open a shared link without an account. Wherever you send it, put a sentence of context and the key steps in the message too, so the recording is a supplement rather than the only way to get the answer; video in customer support has reply templates built that way.
Run this checklist before you send
Most items take seconds once they are habits. Tick them off on the next recording you make; your progress is saved in this browser.
What to do next
Pick the recording your customers see most, usually the top help video or the onboarding walkthrough, and run the checklist on it. Watch it once with the sound off and once with your eyes closed. Whatever you could not follow either way is the fix list: a sentence of narration, a caption correction, a zoomed re-record of one step.
To give a sense of what turns up: on a typical two minute onboarding video, the sound-off pass might find that the captions say "sink your calendar" for "sync your calendar", and that a plan name is wrong in two places. The eyes-closed pass usually finds the silent stretch where the narrator waited for a page to load and then said "there we go" without saying what appeared, and the moment a green badge was the only sign that something worked. Each fix is small. Correct the two caption words, add "the Connected badge appears next to Google Calendar" to the transcript, and next time say it out loud.
If a video fails badly on both passes, re-recording is often quicker than patching. A re-record that follows an outline with locations and results already written in usually takes less time than the original did, and it fixes the captions and transcript at the same time, because they are generated from what you said.
Then change one habit for every new recording: name what you click. It costs nothing, it does the most work of anything on this page, and it makes your captions and transcripts useful by default. When you are ready to go further, how to make tutorial videos puts these habits into a full recording workflow.
Captions and a transcript on every recording VeoRec adds automatic captions and a transcript to every recording on the Free plan, highlights your clicks, and shares a link viewers open without an account. See the Chrome screen recorder
Frequently asked questions
Do screen recordings need captions?
If they are published as web content and have narration, WCAG 2.2 requires captions at Level A, its most basic level. Even where it is not required, captions help anyone watching without sound and anyone more comfortable reading than listening. Automatic captions are a good start but need checking for accuracy.
Are automatic captions good enough for accessibility?
Not on their own. W3C guidance says automatically generated captions do not meet accessibility requirements unless confirmed to be fully accurate. In screen recordings they most often get product names, acronyms and numbers wrong, so watch once with captions on and fix those before publishing.
Does a screen recording need audio description?
Usually not, if you narrate well. WCAG says no additional audio description is needed when all important visual information is already conveyed in the audio. Name each control and its location, say what happens after each action, and read out important values on screen.
What is the difference between captions and a transcript?
Captions are synchronized text shown on the video as it plays. A transcript is the full text of the video on the page, which people can read at their own pace, search and use with a braille display. A descriptive transcript also includes important visual information, which makes it useful to people who are deaf and blind.
What contrast do text overlays in a video need?
WCAG 1.4.3 asks for a contrast ratio of at least 4.5:1 between text and its background, or 3:1 for large text, and it applies to text in images and video too. In practice, put overlay text on a solid dark or light band rather than directly on a busy screenshot.
How do I make a silent screen recording accessible?
A recording with no narration is video-only content, and WCAG 1.2.1 asks for a text alternative or an audio track that presents the same information. The easiest fix is to add narration that describes what happens. Otherwise, publish a text description of every step and result next to the video.