J. Wang, D. Wang, Y. Chen, X. Lu, and C. Zheng, “Noise Robustness Automatic Speech Recognition with Convolutional Neural Network and Time Delay Neural Network,” in Proc. AES Convention 147, Oct. 2019, Paper 10272. [Online]. Available: https://aes.org/publications/elibrary-page/?id=20645
Wang J, Wang D, Chen Y, Lu X, Zheng C. Noise Robustness Automatic Speech Recognition with Convolutional Neural Network and Time Delay Neural Network. In: AES Convention 147. Audio Engineering Society; 2019. Paper 10272. Available from: https://aes.org/publications/elibrary-page/?id=20645
@inproceedings{Wang2019_20645,
author = {Wang, Jie and Wang, Dunze and Chen, Yunda and Lu, Xun and Zheng, Chengshi},
title = {{Noise Robustness Automatic Speech Recognition with Convolutional Neural Network and Time Delay Neural Network}},
booktitle = {AES Convention 147},
note = {Paper 10272},
year = {2019},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=20645}
}
TY - CPAPER
TI - Noise Robustness Automatic Speech Recognition with Convolutional Neural Network and Time Delay Neural Network
AU - Wang, Jie
AU - Wang, Dunze
AU - Chen, Yunda
AU - Lu, Xun
AU - Zheng, Chengshi
T2 - AES Convention 147
M1 - Paper 10272
PY - 2019
DA - 2019/10/06
UR - https://aes.org/publications/elibrary-page/?id=20645
PB - Audio Engineering Society
LA - en
AB - To improve the performance of automatic speech recognition in noisy environments, the convolutional neural network (CNN) combined with time-delay neural network (TDNN) is introduced, which is referred as CNN-TDNN. The CNN-TDNN model is further optimized by factoring the parameter matrix in the time-delay neural network hidden layers and adding a time-restricted self-attention layer after the CNN-TDNN hidden layers. Experimental results show that the optimized CNN-TDNN model has better performance than DNN, CNN, TDNN, and CNN-TDNN. The average recognition word error rate (WER) can be reduced by 11.76% when comparing with the baselines.
ER -