Skip to main content
Category: Assistive Technologies

Speech-to-Text

Also known as: STT, Voice-to-Text, Speech Recognition, Automatic Speech Recognition
Simply put

Speech-to-text is technology that listens to spoken words and converts them into written text. It is commonly used to create transcripts, captions, or dictated notes, and can help people who find typing difficult or who need spoken content presented in a readable form.

Formal definition

Speech-to-text (STT) is a speech recognition process that transcribes spoken language into digital text, often using computational linguistics and, in many current implementations, AI or machine learning models. It is typically delivered as software that processes audio or video input and outputs a text transcript, and it may be described interchangeably as voice-to-text or automatic speech recognition. Accuracy and suitability for a given accessibility use case generally depend on factors such as audio quality, speaker characteristics, and domain-specific vocabulary, so output may require review or correction.

Why it matters

Speech-to-text plays an important role in digital accessibility because it can transform spoken content into a readable form that people can access on their own terms. For individuals who are deaf or hard of hearing, transcripts and captions generated through speech-to-text can make audio and video content perceivable. For people with motor or dexterity limitations who find typing difficult, dictation powered by speech recognition can offer an alternative way to compose text. In this way, the technology supports several distinct groups of users through a single underlying process.

Speech-to-text also intersects with accessibility standards that call for text alternatives to audio and multimedia content. Providing accurate captions and transcripts is commonly cited as a way to help meet recognized guidelines, and speech-to-text can accelerate the production of those alternatives. However, automated output is not automatically compliant or fully accessible. Accuracy generally depends on factors such as audio quality, speaker characteristics, and domain-specific vocabulary, so machine-generated transcripts may contain errors that mislead users or omit meaning. Because of this, output often requires human review and correction before it can be relied upon as an equivalent alternative.

Who it's relevant to

People who are deaf or hard of hearing
Speech-to-text can generate captions and transcripts that make spoken audio and multimedia content perceivable in a readable form. Because automated accuracy varies with audio and speaker conditions, review and correction are often needed to ensure the text faithfully conveys the spoken content.
People with motor or dexterity disabilities
For individuals who find typing difficult, dictation through speech recognition can provide an alternative means of composing notes, messages, and longer documents. The usefulness of dictation depends in part on recognition accuracy and may require editing.
Content creators and media producers
Teams producing video, audio, and other multimedia can use speech-to-text to speed up the creation of transcripts and captions that serve as text alternatives. Producing accurate alternatives is commonly cited as a way to help meet recognized accessibility guidelines, though automated output should be checked before it is treated as an equivalent alternative.
Accessibility and compliance teams
Those responsible for accessibility should understand that speech-to-text is a tool that can support the creation of text alternatives but does not by itself guarantee conformance with WCAG or legal compliance. Manual review and, where appropriate, assistive technology testing remain important. This guidance is not legal advice; consult qualified counsel and current agency rulemaking for obligations specific to your jurisdiction.

Inside STT

Automatic Speech Recognition (ASR)
The underlying technology that converts spoken audio into written text, often powered by machine learning models trained on large speech datasets.
Real-time (live) captioning
Speech-to-text applied to live audio, producing text as speech occurs, commonly used for meetings, events, and live broadcasts. Accuracy and latency vary by tool and speaking conditions.
Transcription of recorded media
Speech-to-text applied to pre-recorded audio or video to generate transcripts and captions, which can be reviewed and corrected before publication.
Human review and correction
The editing step in which people verify and fix automated output for accuracy, speaker identification, punctuation, and specialized terminology that ASR may misrecognize.
Relationship to WCAG captioning and transcript criteria
Speech-to-text output is one means of producing captions and transcripts referenced by WCAG success criteria for time-based media. Meeting those criteria generally depends on accuracy and completeness, not merely the presence of automated text.

Common questions

Answers to the questions practitioners most commonly ask about STT.

Does providing speech-to-text automatically satisfy WCAG requirements for captions?
Not necessarily. Automated speech-to-text output often contains errors in punctuation, speaker identification, and accuracy, particularly with technical terms, accents, or overlapping speakers. WCAG success criteria for captions and transcripts generally require accurate, synchronized text that conveys speech and relevant non-speech information. Raw automated output typically needs human review and editing before it can be relied upon to meet those criteria. Conformance also depends on the specific success criterion and level involved, so speech-to-text should be treated as a starting point rather than a guaranteed solution.
Is speech-to-text the same as captioning?
They are related but not identical. Speech-to-text refers broadly to the technology that converts spoken audio into text and can be used for real-time transcription, captions, or dictation. Captioning specifically involves presenting text synchronized with media, often including speaker identification and non-speech sounds where relevant. Speech-to-text may serve as an input to a captioning workflow, but producing captions that meet accessibility guidelines commonly involves additional editing, formatting, and synchronization work beyond the raw transcription.
When should real-time speech-to-text be used versus a prepared transcript?
Real-time speech-to-text is commonly used for live events, meetings, and broadcasts where content is generated on the fly, and it may be paired with human captioners for greater accuracy. Prepared transcripts are generally appropriate for pre-recorded content, where there is time to review and correct the text. The choice often depends on whether the content is live or recorded, the accuracy required, and the applicable accessibility target for the context.
How can accuracy of speech-to-text output be improved?
Accuracy can often be improved by using high-quality audio input, minimizing background noise, ensuring clear speaker separation, and providing custom vocabularies for specialized terminology. Human review and correction remain important for content intended to meet accessibility guidelines. For live contexts, professional captioners or communication access real-time translation (CART) providers are commonly used to achieve higher accuracy than fully automated systems alone.
What testing is recommended before relying on speech-to-text for accessibility?
Because automated transcription may introduce errors, manual review of the generated text is generally recommended, along with verification that captions are properly synchronized and that speaker changes and relevant non-speech audio are conveyed where appropriate. Testing with the assistive technologies and playback environments your audience uses can help confirm the output functions as intended. Automated tools alone typically detect only a portion of potential issues.
Does speech-to-text address the needs of all users?
Speech-to-text primarily supports users who benefit from a text representation of audio, but it does not by itself address every accessibility need. For example, some users may also require sign language interpretation, and the surrounding interface must still be operable and perceivable through other means. Meeting a technical criterion does not guarantee a fully usable experience for all users, and broader usability testing with people with disabilities is advisable. This entry is informational and not legal advice; requirements evolve through regulation and case law, and qualified counsel should be consulted for specific obligations.

Common misconceptions

Automated speech-to-text is accurate enough to satisfy accessibility requirements on its own.
Automated output commonly contains errors affecting accuracy, punctuation, speaker labeling, and specialized terms. Human review and correction are generally needed for captions and transcripts to convey meaning reliably, and relying solely on unedited automated text may leave content inaccessible.
Providing any speech-to-text output guarantees WCAG conformance or legal compliance.
WCAG success criteria for time-based media generally require accurate and complete captions or transcripts. Producing automated text is a means, not a guarantee of conformance, and conformance itself does not guarantee immunity from legal claims. Consult qualified legal counsel for compliance questions.
Speech-to-text and captioning are the same thing.
Speech-to-text is a technology that produces text from speech; captions and transcripts are specific accessibility deliverables. Captions may also need to include non-speech information and synchronization that raw speech-to-text output does not automatically provide.

Best practices

Treat automated speech-to-text output as a starting draft and apply human review to correct errors, punctuation, and speaker identification before publishing.
For captions, include relevant non-speech information and ensure text is synchronized with the corresponding audio, rather than relying on raw transcript output.
Map deliverables to the applicable WCAG success criteria for time-based media, and confirm the target conformance level (AA is commonly cited) with your accessibility and legal teams.
Verify captions and transcripts with manual and assistive technology testing, since automated tools and automated captions detect and address only a portion of accessibility issues.
Evaluate speech-to-text tools for accuracy under your actual conditions, including accents, technical vocabulary, background noise, and multiple speakers, before adopting them for live use.
Document that any speech-to-text guidance is not legal advice, and consult qualified legal counsel or current agency rulemaking for jurisdiction-specific requirements.