Convention Paper
A Practical Approach to Robust Speech Recognition Using Two Microphones in Driving Environments
AES Convention 137 · Paper 9191 · October 2014
Abstract
Now that the technologies related to the automatic speech recognition have been mature enough and applicable to our everyday life, people have started considering speech as the most desirable human-device interaction means and utilized speech recognition in vehicles. Nonetheless, it is still challenging to recognize speech correctly in driving environments for at least two reasons. One is that the speech signal is corrupted by innumerable noise sources such as the engine sound, road friction, music from the radio, even worse the mixture of spoken words by passengers, etc. Another is that the recognition device may be put at any place like cup holder, passenger seat or dashboard. In this paper we propose a robust speech recognition front-end that removes the probable ambient noise in a driving car regardless of where the recognition device is. The proposed method finds the direction of speech and enhances the speech signal by first detecting the existence of speech utterance using only two microphones. This front-end is designed with practical consideration so that its implementation in the mobile device showed higher recognition accuracy, shorter processing latency and lower computing power consumption than any other top-tier methods.
