Skip to main content
Category: Visual and Media Accessibility

Transcript

Also known as: text transcript, media transcript
Simply put

In general usage, a transcript is a written or typed copy of spoken or recorded material, such as the text of a speech or a recording. In digital accessibility, a transcript is a text version of audio or video content that lets people read what was said and described instead of listening or watching.

Formal definition

A transcript is a written, printed, or typed representation of dictated or recorded material. For accessibility purposes, a transcript provides a text-based alternative to time-based media, typically capturing spoken dialogue and, in a descriptive transcript, relevant non-speech audio and visual information, so that the content is available to users who are deaf, hard of hearing, deafblind (via braille displays), or who otherwise benefit from a readable format. The generic dictionary sense of the term also extends to non-media records, such as an academic transcript documenting a student's courses and grades; the accessibility usage is distinct and refers specifically to text alternatives for audio or video.

Why it matters

For people who are deaf or hard of hearing, a transcript can make the content of audio and video available in a form they can read rather than hear. Transcripts also serve deafblind users who access text through refreshable braille displays, and they benefit anyone who prefers or needs a readable format, including people in sound-sensitive environments or those who find it easier to absorb information by reading. Because a transcript exists as text, it is also searchable and can be indexed, which supports users who want to locate specific information quickly.

It is important to distinguish a transcript from captions. Captions are synchronized with the moving media and appear on screen in time with the audio, while a transcript is a standalone text document covering the same content. A basic transcript typically captures spoken dialogue, whereas a descriptive transcript also includes relevant non-speech audio and important visual information, making it a more complete alternative for video where meaning is conveyed visually. Providing a transcript alone may not be sufficient for all media; for video in particular, users who can see the screen but not hear it are generally better served by captions, and the appropriate combination depends on the content and the standard being applied.

Who it's relevant to

People who are deaf or hard of hearing
Transcripts provide a text version of spoken content, allowing users who cannot hear or who have difficulty hearing audio to read what was said instead of relying on the audio track.
Deafblind users
Because a transcript is text, it can be output to a refreshable braille display, giving deafblind users a way to access media content that is otherwise unavailable to them.
Content creators and media producers
Those who publish audio and video are responsible for generating accurate transcripts, including deciding when a descriptive transcript is needed to convey relevant visual information, and for reviewing automatically generated text before relying on it.
Accessibility and compliance professionals
Teams evaluating digital content need to understand where a transcript is an appropriate text alternative and where captions or additional measures are also required, since providing a transcript alone may not fully address the needs of all users or the requirements of a given standard. This entry is general information and not legal advice; consult qualified counsel and current guidance for specific obligations.
Users who prefer or benefit from reading
People in sound-sensitive settings, non-native language users, and those who process written information more easily can use transcripts to read, search, and reference media content at their own pace.

Inside Transcript

Spoken Dialogue
A text representation of all spoken words in the audio or video content, including narration, conversation, and voiceover.
Speaker Identification
Labels indicating who is speaking, which is particularly important when multiple speakers are present so users can follow the flow of conversation.
Non-Speech Sounds
Descriptions of meaningful non-speech audio, such as [laughter], [applause], or [door slams], that convey information relevant to understanding the content.
Descriptions of Relevant Visual Information
For a descriptive transcript covering video, text descriptions of important visual content that is not otherwise conveyed through the audio, so that users who cannot see the video can access equivalent information.
Accessible Text Format
The transcript is provided as machine-readable text that can be accessed by assistive technologies such as screen readers and refreshable braille displays, and that can be searched or reformatted by the user.

Common questions

Answers to the questions practitioners most commonly ask about Transcript.

Does providing a transcript satisfy accessibility requirements for a video, so captions are not needed?
Generally no. A transcript and captions serve different purposes and are not interchangeable for video. Captions present synchronized text in time with the audio and are important for users who are deaf or hard of hearing to follow content as it plays, while a transcript is a separate text document that may not convey timing. Under WCAG, time-based media typically calls for captions for prerecorded video with audio, and a transcript alone is more commonly associated with audio-only content. For video, a transcript is often a helpful supplement rather than a substitute for captions. This is general guidance and not legal advice; consult current standards and qualified counsel for your specific obligations.
Is an automatically generated transcript sufficient for accessibility?
Not on its own in most cases. Automatically generated transcripts frequently contain errors in wording, speaker identification, punctuation, and terminology, and they may omit relevant non-speech information. Automated tools detect and produce only a portion of what is needed, so human review and correction are generally necessary to produce an accurate transcript. Accuracy matters because an inaccurate transcript can misrepresent the content for the people relying on it. Treat auto-generated output as a starting draft that requires manual verification.
What should a transcript include beyond the spoken words?
A useful transcript commonly includes the spoken dialogue along with speaker identification when more than one person speaks, and descriptions of meaningful non-speech audio such as relevant sounds or audio cues. For content where visual information is essential to understanding, a descriptive transcript may also incorporate descriptions of important visual elements. The goal is to convey the information a user would otherwise miss, so the level of detail should reflect what is meaningful in the source material.
Where should a transcript be placed relative to its media?
A transcript is commonly provided near the associated media, such as directly below the player or via a clearly labeled link adjacent to it, so users can locate it without difficulty. The link or heading should describe what it is, and the transcript itself should be presented as accessible text rather than, for example, an image of text or an inaccessible document. Placement and labeling that make the relationship between the media and its transcript obvious generally support a better experience.
How should transcripts handle multiple speakers and non-speech sounds?
Identify each speaker so readers can follow who is talking, particularly in discussions, interviews, or panels. Non-speech audio that carries meaning, such as relevant sound effects or notable pauses, can be noted in brackets or a comparable convention so the reader understands context that a hearing user would perceive. Consistency in how speakers are labeled and how sounds are described helps readability.
How can the accuracy of a transcript be verified?
Accuracy is generally confirmed through manual review by a person who checks the text against the source audio, correcting wording, speaker attribution, punctuation, and any missing non-speech information. Because automated generation captures only part of what is needed, human proofreading against the original media is the common method for verification. Involving people familiar with the subject matter can help with specialized terminology and proper names.

Common misconceptions

A transcript and captions are the same thing and are interchangeable.
They serve different purposes. Captions are synchronized with the media in time and appear on screen as it plays, while a transcript is a standalone text document of the content. WCAG treats them as distinct, and one does not automatically satisfy the requirement for the other; the appropriate solution depends on whether the content is audio-only, video-only, or synchronized media.
Providing a transcript alone makes any multimedia content fully accessible.
A transcript addresses certain needs but does not on its own satisfy every applicable success criterion. Synchronized video may also require captions and, in some cases, audio description, and meeting individual criteria does not guarantee an accessible experience for all users or immunity from legal claims.
Auto-generated transcripts are sufficient without review.
Automatically generated transcripts frequently contain errors in wording, punctuation, and speaker attribution, and typically omit non-speech sounds and speaker labels. Human review and correction are generally needed to produce an accurate transcript that conveys equivalent information.

Best practices

Determine the media type first: for audio-only content a transcript is commonly the primary text alternative, while for synchronized video you may also need captions and, where applicable, audio description.
Review and correct auto-generated output for accuracy in wording, punctuation, speaker identification, and any non-speech sounds before publishing.
Include clear speaker labels when multiple people are speaking so users can follow the conversation.
For video, consider a descriptive transcript that incorporates meaningful visual information not conveyed through the audio alone.
Provide the transcript as accessible, machine-readable text that works with assistive technologies and can be searched, rather than as an image or inaccessible file.
Place the transcript where users can readily find it in relation to the media, and verify the result with manual and assistive technology testing rather than relying on automated checks alone.