Y. Ke, Y. Hu, J. Li, C. Zheng, and X. Li, “A Generalized Subspace Approach for Multichannel Speech Enhancement Using Machine Learning-Based Speech Presence Probability Estimation,” in Proc. AES Convention 146, Mar. 2019, Paper 10192. [Online]. Available: https://aes.org/publications/elibrary-page/?id=20325
Ke Y, Hu Y, Li J, Zheng C, Li X. A Generalized Subspace Approach for Multichannel Speech Enhancement Using Machine Learning-Based Speech Presence Probability Estimation. In: AES Convention 146. Audio Engineering Society; 2019. Paper 10192. Available from: https://aes.org/publications/elibrary-page/?id=20325
@inproceedings{Ke2019_20325,
author = {Ke, Yuxuan and Hu, Yi and Li, Jian and Zheng, Chengshi and Li, Xiaodong},
title = {{A Generalized Subspace Approach for Multichannel Speech Enhancement Using Machine Learning-Based Speech Presence Probability Estimation}},
booktitle = {AES Convention 146},
note = {Paper 10192},
year = {2019},
month = mar,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=20325}
}
TY - CPAPER
TI - A Generalized Subspace Approach for Multichannel Speech Enhancement Using Machine Learning-Based Speech Presence Probability Estimation
AU - Ke, Yuxuan
AU - Hu, Yi
AU - Li, Jian
AU - Zheng, Chengshi
AU - Li, Xiaodong
T2 - AES Convention 146
M1 - Paper 10192
PY - 2019
DA - 2019/03/06
UR - https://aes.org/publications/elibrary-page/?id=20325
PB - Audio Engineering Society
LA - en
AB - A generalized subspace-based multichannel speech enhancement in frequency domain is proposed by estimating multichannel speech presence probability using machine learning methods. An efficient and low-latency neural networks (NN) model is introduced to discriminatively learn a gain mask for separating the speech and the noise components in noisy scenarios. Besides, a generalized subspace-based approach in frequency domain is proposed, where the speech power spectral density (PSD) matrix and the noise PSD matrix are estimated by short-term and long-term averaging periods, respectively. Experimental results show that the proposed method outperforms the existing NN-based beamforming methods in terms of the perceptual evaluation of speech quality score and the segmental signal-to-noise ratio improvement.
ER -