Skip to main content
Category: Visual and Media Accessibility

Live Captions

Also known as: Live Captioning, Real-Time Captioning, Live Transcription
Simply put

Live captions are text versions of spoken words that appear on screen in real time as someone talks. They are commonly used during live events, lectures, meetings, and video calls to help people who are deaf or hard of hearing, as well as others, follow the audio. Many devices and platforms now include a built-in live captions feature that automatically converts speech to text.

Formal definition

Live captions refer to the real-time conversion of spoken audio into synchronized on-screen text for live presentations, lectures, meetings, and streamed content. Implementations range from operating-system and application-level features that provide automatic speech-to-text transcription (for example, device-level live captions built into iPhone/iOS and Windows, and browser extensions using the device microphone) to professionally produced live captioning workflows for lectures and events. Automatic (machine-generated) live captions vary in accuracy and may not meet the correctness or completeness expectations of human-produced real-time captions; practitioners should evaluate the method used against the intended use case rather than assuming any single approach satisfies a given accessibility requirement. This entry describes the feature category generally and is not legal advice.

Why it matters

Live captions make spoken audio accessible in real time, which is essential for people who are deaf or hard of hearing to follow live lectures, meetings, events, and streamed content as they happen. Because these settings are inherently time-sensitive, captions that appear synchronized with the audio allow participants to engage without waiting for a transcript to be produced afterward. Live captions also assist others, such as people in noisy or sound-sensitive environments and those who process written text more easily.

The method used to generate live captions matters significantly. Automatic (machine-generated) live captions vary in accuracy and may not meet the correctness or completeness expectations of human-produced real-time captioning, particularly where specialized vocabulary, multiple speakers, or poor audio quality are involved. Relying on automatic captions in high-stakes settings, such as academic lectures or public events, may leave gaps that affect a person's ability to fully understand the content.

Because of this variability, practitioners should evaluate the captioning approach against the intended use case rather than assuming any single implementation satisfies a given accessibility need or requirement. Whether a particular method meets applicable legal or institutional obligations depends on the specific context, jurisdiction, and current guidance; this entry describes the feature category generally and is not legal advice. Organizations with questions about their obligations should consult qualified legal counsel.

Who it's relevant to

People who are deaf or hard of hearing
Live captions provide real-time text of spoken audio, helping people who are deaf or hard of hearing follow live lectures, meetings, events, and video calls as they happen.
Educators and academic institutions
Live captioning and transcription are commonly used to add captions to live lectures and events, supporting students who need real-time access to spoken content. The choice between automatic and human-produced captioning should be evaluated against the accuracy needs of the setting.
Event organizers and meeting hosts
Those running live events, conferences, and meetings can use live captions to make spoken content accessible in real time, whether through built-in platform features, browser extensions, or professional live captioning services depending on the accuracy required.
Accessibility practitioners and content teams
Practitioners should assess the captioning method used against the intended use case rather than assuming a single approach meets a given accessibility requirement, weighing the accuracy differences between automatic and human-produced live captions.
General audiences in varied environments
Beyond people who are deaf or hard of hearing, live captions can help anyone better understand audio, including those in noisy or sound-sensitive settings or who prefer to read along with spoken content.

Inside Live Captions

Real-Time Captioning
Text that is generated and displayed synchronously with live audio content, such as during webinars, broadcasts, live streams, or virtual meetings, so that viewers can read spoken dialogue and relevant sounds as they occur.
CART (Communication Access Real-time Translation)
A human-provided captioning method in which a trained stenographer or captioner transcribes speech in real time, generally producing higher accuracy than automated approaches, particularly for specialized vocabulary, multiple speakers, or challenging audio.
Automatic Speech Recognition (ASR)
Software-driven captioning that converts speech to text automatically. It can scale to many events but may produce errors with accents, technical terms, overlapping speakers, or poor audio quality, and typically requires review or supplementation for critical content.
Non-Speech Information
Captions that convey relevant sounds beyond dialogue, such as speaker identification and meaningful audio cues, which help users understand context that is not carried by spoken words alone.
WCAG Relationship
WCAG addresses captions for live audio content through Success Criterion 1.2.4 Captions (Live), which is at Level AA. Meeting this criterion is commonly cited as a target for live multimedia, though conformance alone does not guarantee an accessible experience for all users.
Display and Synchronization
The presentation layer that controls caption timing, readability, positioning, and latency relative to the audio, all of which affect whether users can follow the content in real time.

Common questions

Answers to the questions practitioners most commonly ask about Live Captions.

Do live captions and automatic (AI-generated) captions provide the same level of accuracy?
No. Automatic speech recognition (ASR) can generate captions in real time, but its accuracy varies with audio quality, accents, background noise, technical vocabulary, and speaker overlap. Human-generated live captions, often produced through Communication Access Realtime Translation (CART) by trained captioners, are generally considered more accurate and reliable for meeting accessibility needs. ASR output may be acceptable in some contexts but is commonly viewed as insufficient on its own where accuracy is critical. This guidance is not legal advice; consult qualified counsel for your specific obligations.
Does providing live captions on its own make a live event or stream fully accessible?
Not necessarily. Live captions address one access need, primarily for people who are deaf or hard of hearing, but accessibility for live content may also involve considerations such as sign language interpretation, accessible media players, keyboard operability, and readable caption presentation. Meeting one requirement does not guarantee a fully accessible experience for all users, nor does it guarantee legal compliance. Requirements evolve through regulation and case law, so consult current guidance and qualified legal counsel.
How can caption quality be evaluated during a live event?
Quality can be assessed by monitoring factors such as accuracy of the transcribed words, latency between speech and displayed text, correct speaker identification where relevant, and readability of the caption presentation. Because live captioning happens in real time, organizers commonly plan for monitoring and, where feasible, a way to correct or supplement captions. Manual review and testing with assistive technology and actual users can identify issues that automated checks may miss.
What can be done to improve accuracy when using automatic captions?
Common practices include providing clear, high-quality audio, minimizing background noise, having speakers talk at a measured pace and avoid talking over one another, and supplying the ASR system with specialized terminology or names in advance where the tool supports it. Even with these steps, output accuracy may still vary, so organizations often consider human captioning for higher-stakes or accuracy-sensitive contexts.
How should latency between spoken words and displayed captions be handled?
Some delay is inherent to live captioning, whether produced by humans or ASR, because the audio must be processed and rendered as text. Planning for this involves setting realistic expectations, choosing tools and services that keep latency manageable, and, for scripted portions, considering prepared text where appropriate. Testing under conditions similar to the live event can help identify whether latency is acceptable for the audience.
Where should live captions be displayed so they are usable?
Captions should generally be presented so that viewers can read them without losing access to the main content, whether within the media player, in a synchronized side panel, or through an integrated captioning display, depending on the platform. Readability considerations commonly include adequate contrast, sufficient text size, and, where possible, user control over caption appearance. Testing with the intended platforms and with assistive technology helps confirm the display works for the audience.

Common misconceptions

Automated live captions are accurate enough to meet accessibility needs on their own.
Automatic speech recognition can vary in accuracy depending on audio quality, accents, terminology, and multiple speakers. For high-stakes or complex content, human CART captioning is often preferred, and automated output may need monitoring or correction. Accuracy limitations mean automated captions may not fully serve all users.
Live captions only need to include spoken words.
To be useful, captions generally should identify speakers and convey meaningful non-speech audio where relevant, so that users who rely on captions receive comparable context to those hearing the audio.
Providing live captions guarantees compliance with the law.
Captions can support conformance with WCAG Success Criterion 1.2.4 (Level AA), but meeting a technical criterion does not by itself guarantee legal compliance or an accessible experience. Legal obligations depend on the applicable authority and evolving case law and regulation; consult qualified legal counsel for specific situations.

Best practices

Determine whether human CART captioning or automatic speech recognition is appropriate for the event, choosing human captioners for high-stakes, specialized, or complex multi-speaker content where accuracy is critical.
Include speaker identification and meaningful non-speech audio information so captions convey context comparable to the audio.
Monitor caption quality, latency, and synchronization during live events, and have a process to correct or supplement inaccurate automated output.
Target WCAG 2.1/2.2 Success Criterion 1.2.4 Captions (Live) at Level AA as a commonly cited benchmark for live audio content, while recognizing that conformance alone does not ensure usability for all users.
Test the caption display for readability, positioning, and reliable delivery with real users and assistive technologies rather than relying solely on automated checks.
Confirm the accessibility of the delivery platform's captioning features and provide a way for participants to request accommodations in advance.