Skip to main content
Category: Visual and Media Accessibility

Closed Captions

Also known as: CC, Closed Captioning, CC
Simply put

Closed captions are text versions of the audio in a video that you can turn on or off, showing what is being said along with relevant sounds. They let people who are deaf or hard of hearing follow along with visual content, and can also help viewers in noisy or quiet environments. Because they are 'closed,' the viewer chooses whether to display them on screen.

Formal definition

Closed captions are time-synchronized text tracks that render an audio program's spoken dialogue and, unlike subtitles intended primarily for language translation, are generally designed to also convey non-speech audio information and speaker context to support viewers who are deaf or hard of hearing. They are termed 'closed' because they can be toggled on or off by the user, in contrast to open captions that are permanently embedded in the video image. Closed captions are commonly distinguished from subtitles, which typically focus on translating dialogue into another language and may assume the viewer can hear other audio cues.

Why it matters

Closed captions are a core component of media accessibility because they give people who are deaf or hard of hearing access to spoken dialogue and other meaningful audio content in video. Without captions, a video's information may be entirely inaccessible to these viewers, which is why captioning is frequently addressed in accessibility standards and is often expected for audio and video content published by organizations. Because closed captions can be toggled on or off, they preserve viewer choice while still making the content available to those who need it.

Captions also benefit a broader audience beyond those they are primarily designed for. Viewers in noisy environments, such as public spaces, or in quiet settings where audio cannot be played, can follow content through captions. This broad utility is one reason captioning is commonly treated as a baseline practice rather than an optional enhancement for video content.

It is worth distinguishing closed captions from subtitles. Subtitles typically focus on translating dialogue into another language and may assume the viewer can hear other audio cues, while closed captions are generally designed to also convey non-speech audio information and speaker context to support viewers who are deaf or hard of hearing. Confusing the two can result in captions that omit important sounds, so understanding this distinction matters when producing accessible media. This entry is general guidance and not legal advice; specific obligations depend on jurisdiction, applicable regulation, and evolving case law, and organizations should consult qualified counsel where legal questions arise.

Who it's relevant to

People who are deaf or hard of hearing
Closed captions are primarily intended to aid individuals who are deaf or hard of hearing by presenting the audio portion of a program as on-screen text. For these viewers, captions can be the difference between fully understanding a video and being excluded from its content.
Content creators and video producers
Those who produce and publish video are responsible for providing captions that accurately and completely represent the audio, including non-speech information where relevant. Understanding the difference between captions and subtitles helps producers avoid omitting audio cues that deaf and hard-of-hearing viewers rely on.
Accessibility and compliance teams
Teams responsible for accessibility often treat captioning of video content as a baseline expectation and evaluate whether captions are present, accurate, and synchronized. Because requirements vary by jurisdiction and evolve through regulation and case law, these teams should confirm specific obligations with qualified legal counsel rather than assuming a single universal rule.
General viewers in situational contexts
Viewers who cannot or prefer not to use audio, such as those in noisy public spaces or quiet environments, can benefit from the ability to turn captions on. This situational usefulness is one reason captions are widely adopted beyond their primary accessibility purpose.

Inside CC

Synchronized Text
Text that appears on screen in time with the corresponding audio, conveying spoken dialogue as it occurs so users who cannot hear the audio can follow along.
Non-Speech Audio Information
Descriptions of relevant sounds beyond dialogue, such as music, sound effects, laughter, or off-screen noises, that carry meaning for understanding the content.
Speaker Identification
Labels or cues indicating who is speaking when this is not otherwise clear from the video, helping users attribute dialogue to the correct person.
User Toggle Control
The defining feature that distinguishes closed captions from open captions: closed captions can be turned on or off by the viewer, whereas open captions are permanently embedded in the video.
Relationship to WCAG
Captions for prerecorded synchronized media are addressed by WCAG success criterion 1.2.2 (Captions, Prerecorded) at Level A, and captions for live synchronized media are addressed by success criterion 1.2.4 (Captions, Live) at Level AA. These criteria are commonly cited benchmarks for accessible media.

Common questions

Answers to the questions practitioners most commonly ask about CC.

Are closed captions the same as subtitles?
Not quite. The terms are often used interchangeably, but they generally serve different purposes. Subtitles typically assume the viewer can hear the audio and primarily translate or transcribe spoken dialogue, often for language access. Closed captions are designed for viewers who are deaf or hard of hearing and include not only dialogue but also relevant non-speech audio information such as speaker identification, sound effects, and music cues. Because of this broader scope, closed captions are more commonly associated with accessibility requirements, though the exact expectations can vary by context and applicable standards.
Do automatically generated captions satisfy accessibility requirements?
Not on their own, in most cases. Automatically generated captions can be a useful starting point, but they frequently contain errors in wording, punctuation, speaker identification, and non-speech audio, and they may not accurately reflect the content. Captions are generally expected to be accurate and synchronized with the audio to be considered accessible. Auto-generated output typically requires human review and correction before it can reasonably meet common accessibility expectations. Automated tools also do not replace the need for manual review, and conformance with a standard does not by itself guarantee legal compliance. For questions about specific obligations, consult qualified legal counsel.
What is the difference between closed captions and open captions?
Closed captions can be turned on or off by the viewer, giving users control over whether the text is displayed. Open captions are permanently burned into the video and cannot be disabled. Closed captions are generally preferred where the delivery platform supports them because they offer user control and can often be styled or repositioned, but open captions may be used when the playback environment does not reliably support a caption track.
What should closed captions include beyond spoken dialogue?
To support viewers who are deaf or hard of hearing, closed captions should generally convey meaningful audio information beyond the words themselves. This commonly includes speaker identification when it is not otherwise clear, relevant sound effects, and indications of music or other significant non-speech audio. The goal is to give caption users access to the same essential information that is available through the audio track.
How are captions typically added to web video?
Captions are commonly delivered as a separate timed-text file associated with the video rather than being embedded directly into the picture. In web contexts, caption tracks are often provided through supported caption or timed-text formats and referenced by the video player so the viewer can enable or disable them. Because support and behavior can vary across players and platforms, it is advisable to verify that captions display correctly and remain synchronized in the environments where the content will be viewed.
How should closed captions be tested for quality?
Testing generally involves reviewing captions for accuracy of wording, correct timing and synchronization with the audio, appropriate speaker identification, and inclusion of relevant non-speech audio. Because automated checks detect only a portion of potential issues, manual review is important, ideally including playback with the captions enabled to confirm they display correctly, are readable, and reflect the content. Involving users of captions or accessibility reviewers can help identify problems that automated tools may miss.

Common misconceptions

Closed captions and subtitles are the same thing.
Subtitles generally assume the viewer can hear the audio and typically render only spoken dialogue, often for translation purposes. Closed captions are intended for users who cannot hear the audio and additionally convey non-speech information such as sound effects, music, and speaker identification.
Automatically generated captions are sufficient for accessibility and compliance.
Automated speech recognition can contain errors in wording, punctuation, speaker attribution, and timing, and may omit non-speech audio. Human review and correction are commonly needed for captions to accurately and fully convey the content, and accuracy is central to their usefulness.
Adding closed captions guarantees WCAG conformance or legal compliance.
Captions address specific success criteria but do not by themselves ensure overall conformance or immunity from legal claims. Other requirements, such as audio description for visual content, may also apply, and legal exposure depends on jurisdiction, regulation, and case law. This guidance is not legal advice.

Best practices

Provide captions that include not only dialogue but also relevant non-speech audio such as sound effects and music, plus speaker identification where the speaker is not otherwise clear.
Treat automated captions as a starting point and have a person review and correct wording, punctuation, timing, and attribution for accuracy.
Synchronize caption text closely with the corresponding audio so viewers can follow along without losing context.
Offer captions as user-controllable (closed) so viewers can turn them on or off, and ensure the toggle is discoverable in the media player.
Target WCAG 1.2.2 (Level A) for prerecorded media and 1.2.4 (Level AA) for live media as commonly cited benchmarks, and consider related requirements such as audio description where applicable.
Verify captions through manual review and, where relevant, with assistive technology, since automated checks alone detect only a portion of caption quality issues.