L. Cheng, C. Zheng, R. Peng, and X. Li, “Improvement of DNN-Based Speech Enhancement with Non-Normalized Features by Using an Automatic Gain Control,” in Proc. AES Convention 147, Oct. 2019, Paper 10256. [Online]. Available: https://aes.org/publications/elibrary-page/?id=20629
Cheng L, Zheng C, Peng R, Li X. Improvement of DNN-Based Speech Enhancement with Non-Normalized Features by Using an Automatic Gain Control. In: AES Convention 147. Audio Engineering Society; 2019. Paper 10256. Available from: https://aes.org/publications/elibrary-page/?id=20629
@inproceedings{Cheng2019_20629,
author = {Cheng, Linjuan and Zheng, Chengshi and Peng, Renhua and Li, Xiaodong},
title = {{Improvement of DNN-Based Speech Enhancement with Non-Normalized Features by Using an Automatic Gain Control}},
booktitle = {AES Convention 147},
note = {Paper 10256},
year = {2019},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=20629}
}
TY - CPAPER
TI - Improvement of DNN-Based Speech Enhancement with Non-Normalized Features by Using an Automatic Gain Control
AU - Cheng, Linjuan
AU - Zheng, Chengshi
AU - Peng, Renhua
AU - Li, Xiaodong
T2 - AES Convention 147
M1 - Paper 10256
PY - 2019
DA - 2019/10/06
UR - https://aes.org/publications/elibrary-page/?id=20629
PB - Audio Engineering Society
LA - en
AB - Speech enhancement performance may degrade when the peak level of the noisy speech is significantly different from the training datasets in Deep Neural Networks (DNN)-based speech enhancement algorithms, especially when the non-normalized features are used in practical applications, such as log-power spectra. To overcome this shortcoming, we introduce an automatic gain control (AGC) method as a preprocessing technique. By doing so, we can train the model with the same peak level of all the speech utterances. To further improve the proposed DNN-based algorithm, the feature compensation method is combined with the AGC method. Experimental results indicate that the proposed algorithm can maintain consistent performance when the peak of the noisy speech changes in a large range.
ER -