Skip to content
AI for Beginners

Speech-to-Text(STT)

AI that listens to spoken words and writes them down as text. It is what turns a voice note into a message, powers the dictation button on your phone, and produces the captions on a video. This is called speech-to-text, sometimes shortened to STT, and it is also known as speech recognition or transcription.

Speech-to-text is the technology that converts talking into written words. You speak, and it types. It sits behind a great many everyday features: dictating a text message, the automatic subtitles on a video, voice assistants, and tools that transcribe a meeting or an interview into a document you can read and search.

Modern speech-to-text is impressively good, though it is not flawless. It can stumble over strong accents, background noise, technical jargon, unusual names, and people talking over one another. Knowing this helps you use it well. A quick read-through of anything it produces will usually catch the odd word it misheard.

For many people this is one of the most quietly useful forms of AI, and it costs nothing to try. Speaking is often faster than typing, so dictating a first draft and tidying it up afterwards can be a real time-saver, especially on a phone or when your hands are busy. It is also a genuine accessibility aid for anyone who finds typing difficult.

The reverse also exists and is worth a mention: text-to-speech turns written words into a spoken voice, which is how screen readers and audiobook-style tools work. Speech-to-text handles the listening; text-to-speech handles the talking.

Related terms