Speech recognition, also called automatic speech recognition or speech-to-text, is technology that translates spoken language into text. Researchers Stephen Balashek, R. Biddulph and K. H. Davis at Bell Labs built Audrey in 1952, the first system able to recognize spoken digits from a single speaker by locating formants in the power spectrum of an utterance, and IBM demonstrated a sixteen-word recognition machine nicknamed Shoebox at the 1962 World's Fair. Raj Reddy, as a graduate student at Stanford in the late 1960s, became the first person to work on continuous speech recognition, removing the earlier requirement that a speaker pause between words, and James Baker and Janet Baker at Carnegie Mellon began applying hidden Markov models in the 1970s to combine acoustic, language and syntactic information into a single probabilistic system. Fred Jelinek's team at IBM built on statistical approaches like these to create Tangora in the mid-1980s, a voice-activated typewriter able to handle a vocabulary of twenty thousand words.
Facts
SignificanceTranslates spoken language into text or other interpretable forms, underpinning voice interfaces, dictation and transcription. 1 Sources
1. Speech recognition, Wikipedia
History, 1952
Bell Labs researchers, Stephen Balashek, R. Biddulph, and K. H. Davis, built Audrey for single-speaker digit recognition.
Lead section, paragraph 1
is a sub-field of computational linguistics concerned with methods and technologies that translate spoken language into text or other interpretable forms.
View the SourceReader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.