Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Noise Robustness Automatic Speech Recognition with Convolutional Neural Network and Time Delay Neural Network

Authors: Wang, Jie; Wang, Dunze; Chen, Yunda; Lu, Xun; Zheng, Chengshi

AES Convention 147 · Paper 10272 · October 2019

Abstract

To improve the performance of automatic speech recognition in noisy environments, the convolutional neural network (CNN) combined with time-delay neural network (TDNN) is introduced, which is referred as CNN-TDNN. The CNN-TDNN model is further optimized by factoring the parameter matrix in the time-delay neural network hidden layers and adding a time-restricted self-attention layer after the CNN-TDNN hidden layers. Experimental results show that the optimized CNN-TDNN model has better performance than DNN, CNN, TDNN, and CNN-TDNN. The average recognition word error rate (WER) can be reduced by 11.76% when comparing with the baselines.

Details

Published in
AES Convention 147
AES Convention
147
Paper number
10272
Publication date
October 6, 2019
Session subject
Posters: Applications in Audio
Affiliation
Guangzhou University, Guangzhou, China; Power Grid Planning Center, Guandgong Power Grid Company, Guangdong, China; Institute of Acoustics, Chinese Academy of Sciences, Beijing, China (See document for exact affiliation information.)
Type
Convention Paper