Speech-to-Text
Speech-to-text is technology that listens to spoken words and converts them into written text. It is commonly used to create transcripts, captions, or dictated notes, and can help people who find typing difficult or who need spoken content presented in a readable form.
Speech-to-text (STT) is a speech recognition process that transcribes spoken language into digital text, often using computational linguistics and, in many current implementations, AI or machine learning models. It is typically delivered as software that processes audio or video input and outputs a text transcript, and it may be described interchangeably as voice-to-text or automatic speech recognition. Accuracy and suitability for a given accessibility use case generally depend on factors such as audio quality, speaker characteristics, and domain-specific vocabulary, so output may require review or correction.
Why it matters
Speech-to-text plays an important role in digital accessibility because it can transform spoken content into a readable form that people can access on their own terms. For individuals who are deaf or hard of hearing, transcripts and captions generated through speech-to-text can make audio and video content perceivable. For people with motor or dexterity limitations who find typing difficult, dictation powered by speech recognition can offer an alternative way to compose text. In this way, the technology supports several distinct groups of users through a single underlying process.
Speech-to-text also intersects with accessibility standards that call for text alternatives to audio and multimedia content. Providing accurate captions and transcripts is commonly cited as a way to help meet recognized guidelines, and speech-to-text can accelerate the production of those alternatives. However, automated output is not automatically compliant or fully accessible. Accuracy generally depends on factors such as audio quality, speaker characteristics, and domain-specific vocabulary, so machine-generated transcripts may contain errors that mislead users or omit meaning. Because of this, output often requires human review and correction before it can be relied upon as an equivalent alternative.
Who it's relevant to
Inside STT
Common questions
Answers to the questions practitioners most commonly ask about STT.