Skip to main content
Category: Visual and Media Accessibility

Captions

Also known as: Closed Captions, Open Captions, Captioning
Simply put

Captions are lines of text displayed on a video or other media that transcribe the spoken words and identify other important sounds, such as music or sound effects. They allow people who are deaf or hard of hearing, as well as viewers in sound-off environments, to follow along with audio content. Captioning is the process of converting a program's audio into this synchronized text.

Formal definition

Captions are a time-synchronized text alternative for the audio track of media content, presenting spoken dialogue along with non-speech audio information (for example, speaker identification, sound effects, and music) needed to understand the content. Captioning is the process of converting audio content from broadcasts, webcasts, films, video, live events, and similar productions into this text. Captions are commonly distinguished as 'closed' (able to be turned on or off by the viewer) or 'open' (permanently rendered into the video). Captions differ from subtitles, which traditionally convey only dialogue translation and assume the viewer can hear other audio. Note that captions are one component of media accessibility and support conformance-related goals for time-based media; a full technical treatment of applicable success criteria and conformance levels is outside the scope of this core definition.

Why it matters

Captions provide access to audio content for people who are deaf or hard of hearing, allowing them to follow spoken dialogue and understand important non-speech sounds such as music, sound effects, and speaker identification. Without captions, video and other time-based media can exclude a significant portion of the audience from information, entertainment, education, and services delivered through audio. Captions also benefit people who are not disabled, including viewers in sound-off environments such as public spaces, offices, or transit, and those watching content in a language they are still learning.

In the context of digital accessibility, captions are one of the primary components of accessible time-based media and are commonly referenced when organizations work toward WCAG-related goals for video content. Providing captions supports the goal of offering a text alternative to audio so that content is perceivable to users who cannot rely on sound. However, captions alone do not address every media accessibility need; for example, they do not convey visual information to users who cannot see the screen, which is handled through other techniques such as audio description.

Because the specific obligations to caption content can depend on the applicable legal framework, jurisdiction, and setting, organizations should treat captioning as part of a broader accessibility strategy rather than a standalone guarantee of compliance. Requirements evolve through regulation and case law, and this entry is not legal advice; readers with specific obligations should consult qualified legal counsel or current agency guidance.

Who it's relevant to

Content and media producers
Teams that create video, webcasts, films, and live events are responsible for producing captions that accurately transcribe dialogue and convey relevant non-speech audio. They decide between open and closed captions and establish workflows, including review of any automatically generated text, to ensure captions are synchronized and complete.
People who are deaf or hard of hearing
Captions are a primary means of access to spoken dialogue and important sounds for viewers who cannot rely on audio. Well-produced captions allow these users to follow along with the content, while inaccurate or missing captions can exclude them from the information being conveyed.
Accessibility and UX practitioners
Engineers, designers, and accessibility specialists incorporate captioning into media workflows and testing as part of supporting accessible time-based media. They should be aware that captions are one component of media accessibility and that conformance details for time-based media are addressed by specific success criteria beyond this core definition.
Compliance officers and legal counsel
Those responsible for accessibility obligations evaluate where captioning applies within their organization's content and setting. Because specific requirements depend on the applicable legal framework and jurisdiction and continue to evolve, this audience should rely on current agency guidance and qualified legal counsel rather than treating captions as a standalone guarantee of compliance.
Viewers in sound-off environments
People watching video without audio, for example in public spaces, offices, or on transit, benefit from captions as a way to follow content when sound is unavailable or impractical, illustrating the broader usability value of captioning beyond disability access.

Inside Captions

Synchronized text
Captions present the audio content as text that is timed to appear in sync with the corresponding speech, sounds, and events in the media.
Speech transcription
The spoken dialogue and narration in the media, transcribed into readable text and attributed to speakers where relevant to comprehension.
Non-speech audio information
Descriptions of meaningful sounds beyond dialogue, such as music, laughter, applause, or sound effects, which distinguish captions from a plain dialogue transcript.
Closed versus open captions
Closed captions can be turned on or off by the user, while open captions are permanently embedded in the video image and cannot be disabled.
Relationship to WCAG success criteria
Captioning is addressed by WCAG success criteria concerning captions for prerecorded and live time-based media, with prerecorded captions commonly cited at Level A and live captions generally associated with Level AA.

Common questions

Answers to the questions practitioners most commonly ask about Captions.

Are captions and subtitles the same thing?
Not exactly. The terms are often used interchangeably, but they are generally distinguished by purpose. Captions are intended for viewers who cannot hear the audio and typically include not only dialogue but also relevant non-speech information such as speaker identification and important sound effects. Subtitles traditionally assume the viewer can hear the audio and often provide only a translation or transcription of spoken dialogue. For accessibility purposes, the goal is to convey the full auditory experience, which is why captions, rather than dialogue-only subtitles, are commonly referenced in relation to WCAG success criteria.
Do automatically generated captions satisfy accessibility requirements?
Not on their own, in most cases. Auto-generated captions can serve as a starting point, but they frequently contain errors in wording, punctuation, speaker identification, and timing, and they may omit non-speech audio information. Because inaccurate captions can misrepresent content, they generally do not by themselves meet the intent of the relevant WCAG success criteria, which call for captions that accurately convey the audio. Human review and correction are commonly recommended. Note that meeting a technical criterion does not by itself guarantee a fully usable experience or legal compliance.
What is the difference between closed and open captions?
Closed captions can be turned on or off by the viewer and are delivered separately from the video image, often via a caption track or file. Open captions are permanently embedded into the video image and cannot be disabled. Closed captions are commonly preferred because they give users control and can support multiple languages or formats, while open captions may be used where a player does not reliably support caption tracks. The appropriate choice may depend on your delivery platform and audience needs.
Which WCAG success criteria are most relevant to captions?
Captions relate to success criteria addressing time-based media. Criteria at Level A generally address captions for prerecorded content, while additional criteria at Level AA commonly address captions for live content. Level AA is the conformance level most frequently cited as a target. Because criteria and their level assignments can differ across WCAG 2.0, 2.1, and 2.2, you should consult the specific version you are targeting to confirm which criteria apply and at what level.
What information should captions include beyond spoken dialogue?
To accurately convey the audio, captions generally should include speaker identification when it is not otherwise clear, and relevant non-speech sounds such as meaningful sound effects, music cues, or other audio that affects understanding of the content. The aim is for a viewer who cannot hear the audio to receive equivalent information to someone who can. Judgment is often required to determine which sounds are meaningful in context.
How should captions be tested for accuracy and quality?
Testing commonly involves manual review rather than reliance on automated tools alone. Reviewers typically check that captions match the spoken words, are correctly timed and synchronized with the audio, are readable, identify speakers where needed, and include relevant non-speech audio. Testing with the actual media player and with assistive technology can help confirm that caption controls function as expected. Automated methods can detect only a portion of potential issues, so manual verification is generally necessary.

Common misconceptions

Captions and subtitles are the same thing.
Subtitles generally assume the viewer can hear and typically convey only dialogue, often for translation, whereas captions are intended for users who cannot hear the audio and include non-speech sounds and speaker identification where needed for comprehension.
Automatically generated captions are sufficient for accessibility and compliance.
Auto-generated captions frequently contain errors in wording, punctuation, speaker identification, and timing, and often omit non-speech audio. They generally require human review and correction to accurately reflect the audio content.
Adding captions alone makes multimedia fully accessible.
Captions primarily serve users who are deaf or hard of hearing. Full media accessibility may also require other measures, such as audio description for visual content and accessible media player controls, and meeting a single criterion does not guarantee an accessible experience for all users or immunity from legal claims.

Best practices

Provide captions for prerecorded audio-visual content, and provide real-time captions for live time-based media where applicable, consistent with the relevant WCAG success criteria.
Include non-speech audio information, such as relevant sound effects and music, and identify speakers when it aids comprehension, rather than transcribing dialogue alone.
Review and correct any automatically generated captions to ensure accuracy in wording, punctuation, timing, and synchronization before publishing.
Offer closed captions that users can enable or disable, and ensure the media player exposes accessible controls for toggling them.
Verify caption quality through manual review and testing with assistive technology, since automated checks detect only a portion of accessibility issues.
Treat captions as one component of media accessibility and consider complementary measures such as transcripts and audio description; consult qualified legal counsel and current agency guidance regarding jurisdiction-specific requirements, which evolve through regulation and case law.