Automatic Speech Recognition: Research Progress, Achievements, and Challenges
On Monday, March 5, 2012, Prof. Sadaoki Furui (professor emeritus of Tokyo Institute of Technology) gave a public lecture to students taking the IF3054 Artificial Intelligence and IF6058 Natural Language Processing courses in the Informatics Engineering Study Program at R.7602 Benny Subianto Building (Labtek V) on human speech recognition technology. Automated speech recognition technology (automatic speech recognition (ASR)is a series of technologies capable of recognizing human speech and converting it into a series of texts. Prof. Sadaoki Furui himself is one of the world's scientists who is widely involved in various research in the fields of speech analysis, voice/speech recognition, voice identity recognition, and speech synthesis, the results of his research have been published in more than 350 international journals.
In the lecture, Prof. Sadaoki Furui outlined the development of speech recognition technology over the 30 years since the research began, along with the achievements, challenges, and obstacles encountered. Broadly speaking, research developments in automated speech recognition can be divided into four generations, each influenced by developments in information technology at the time.
The first generation (1G) of research in this field took place from 1952 to 1970. Researchers in this generation (Bell labs, RCA Labs, MIT Lincoln Labs (Amerika Serikat), University College London (English), and Radio Research Lab, Kyoto Univ., NEC Labs (Japanese)) attempts to recognize digits/syllables/vowel sounds/phonemes using a heuristic approach.
In the early 1970s, the approach used shifted towards pattern adjustment (pattern matching) which marked the beginning of the second generation (2G). In this generation, the DTW technique was known (dynamic time warping) used by Vintsyuk (Russia) and NEC labs (Japan). Speech recognizers that can recognize isolated words have begun to be used in various applications, especially in Russia and Japan. In addition, speech recognition systems that can recognize large numbers of words (large-vocabulary ASR system) began to be developed at IBM Labs and a speech recognition system capable of recognizing all users (speaker-independent ASR) also began to be developed at Bell Labs. Another important development in this generation was the development of a speech recognition system that was able to recognize continuous speech (continuous speech recognition) by using dynamic phoneme search performed by Carnegie Mellon UniversityTechnology to not only recognize speech but also understand the meaning of that speech is also starting to be developed through the DARPA program.
The beginning of the third generation (3G) of speech recognition technology began in 1980, marked by the use of a statistical approach. In speech recognition systems using a statistical approach, Hidden Markovs Model (HMM) used to create acoustic models, while for modeling language it is starting to be used n-gram. For sound processing used cepstrum + deltacepstrum. In addition, in the third generation, speech recognition systems based on artificial neural networks (neural network). Automated speech recognition-based resource management applications were developed through DARPA programs at various research institutions/universities worldwide such as the SPHINX system (CMU), BYBLOS (BBN), DECIPHER (SRI), and at Lincoln Labs, MIT, AT&T Bell Labs.
To date, statistical approaches using HMMs and n-grams are still used because they have been proven to produce good recognition results. To improve the performance of statistical-based speech recognition systems, researchers at various institutions worldwide have developed various techniques, such as Error minimization (discriminative) approach, VTLN, MLLR, HLDA, fMPE, PMC to reduce noise caused by the surrounding environment, individual characteristics, microphones, transmission channels, and others. This phase, which began in the early 1990s, is known as the 3.5G generation because it is an improvement over the 3G phase. In the 2000s, spontaneous speech recognition (spontaneous speech recognition) began to be developed in Japan with the construction of a national-scale spontaneous conversation corpus (CSJ corpus), and in America and Europe through Meeting projects. In addition, this generation introduced multimodal speech recognition, which combines audio and visual input to enhance the system's capabilities.
Currently, research in the field of speech recognition has advanced to the fourth generation (4G). The direction of fourth-generation research is toward speech understanding, namely developing techniques that enable computers to understand human speech, not just recognize or convert human speech into text as in previous generations. This requires a deeper understanding of human speech processing. Significant progress can be achieved by using data-intensive approaches to extract knowledge. active learning, unsupervised, semi-supervised or lightly-supervised training/adaptation become a very important technology at this time.
Over the past 30 years, research in the field of speech recognition has developed very rapidly. The most successful applications developed include applications that enable humans and computers to conduct interactive dialogue using speech (spoken dialogand automatic transcription. However, this technology still has several weaknesses, including handling voice variations caused by individual characteristics, noise, language/dialect variations, and conversation topics. Another problem that remains unaddressed is how to recognize words that are not in the system's dictionary, known as out-of-vocabulary (OOV). For researchers who want to develop a speech recognition system using a new language, building such a system takes a considerable amount of time. The biggest obstacle lies in the unavailability of the data needed to build such a system. This data includes large-scale speech data (voice corpus) and text data (text corpus).
At the end of the lecture, participants interactively asked various questions. The highly engaging lecture concluded at 12:00 PM. Prof. Sadaoki Furui was very impressed with the questions raised by the participants and hoped that more Indonesian researchers would be interested in continuing research in the field of speech recognition technology. One of the major tasks fundamental to the development of this research is the availability of a large-scale speech and text corpus that can represent all the diversity that exists in the Indonesian language.
Written By: Dessi Puji Lestari